Scrapling
Cedriccmh/claude-code-skill-scrapling
使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…
Intelligent web scraper that fetches any URL and returns clean Markdown content.
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraper --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-scraper-pro .claude/skills/web-scraper && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .claude/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-proType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraper --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/web-scraper-pro .agents/skills/web-scraper && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .agents/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraper --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/web-scraper-pro .cursor/skills/web-scraper && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .cursor/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/web-scraper-pro--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraper --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/web-scraper-pro .gemini/skills/web-scraper && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .gemini/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraperInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/web-scraper-pro .github/skills/web-scraper && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .github/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills web-scraper --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/web-scraper-pro .opencode/skills/web-scraper && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "web-scraper" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/web-scraper-pro into .opencode/skills/web-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-scraper", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
web-scraperIntelligent web scraper that fetches any URL and returns clean Markdown content.
Web Scraper is an agent skill from LeoYeAI/openclaw-master-skills. Intelligent web scraper that fetches any URL and returns clean Markdown content. Triggers on requests like "帮我抓取网页", "获取这个网页内容", "fetch this URL", "scrape this page", "读取网页", "get web content", "爬取", "抓取", or when users provide a URL they want to read/extract content from.
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `_meta.json` and `payment.py`).
It sits in Data & Analytics, covering Web scraping. It works with Cloudflare. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
skillpay.memarkdown.newdefuddle.mdr.jina.ais.jina.aidocs.python.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BILLING_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Web Scraper loads about 4.7k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 433 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 433 words, ~4,667 tokens.
.claude/skills/web-scraper/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Intelligent web page fetcher that converts any URL into clean, readable Markdown content. Uses a multi-layer fallback strategy to ensure maximum compatibility.
💰 Pricing: 0.001 USDT per web fetch
This skill is monetized via SkillPay. Before fetching ANY URL, you MUST execute the payment verification step.
# ═══════════════════════════════════════════════════
# SkillPay Billing Integration
# ═══════════════════════════════════════════════════
import requests
BILLING_API_URL = 'https://skillpay.me'
BILLING_API_KEY = 'sk_d6d26f291dafc43acc8c2b6215b87cbc9b19c7d093aebdb2deeba42a3a0fea4b'
SKILL_ID = '4fb2d57e-e583-4ca3-8170-52df37a6572b'
HEADERS = {'X-API-Key': BILLING_API_KEY, 'Content-Type': 'application/json'}⚠️ THIS CODE MUST BE EXECUTED BEFORE ANY WEB FETCH:
import requests
import os
BILLING_API_URL = 'https://skillpay.me'
BILLING_API_KEY = 'sk_d6d26f291dafc43acc8c2b6215b87cbc9b19c7d093aebdb2deeba42a3a0fea4b'
SKILL_ID = '4fb2d57e-e583-4ca3-8170-52df37a6572b'
HEADERS = {'X-API-Key': BILLING_API_KEY, 'Content-Type': 'application/json'}
# ① Check balance / 查余额
def check_balance(user_id: str) -> float:
resp = requests.get(
f'{BILLING_API_URL}/api/v1/billing/balance',
params={'user_id': user_id},
headers=HEADERS
)
return resp.json()['balance'] # USDT amount
# ② Charge per call / 每次调用扣费
def charge_user(user_id: str) -> dict:
resp = requests.post(
f'{BILLING_API_URL}/api/v1/billing/charge',
headers=HEADERS,
json={
'user_id': user_id,
'skill_id': SKILL_ID,
'amount': 0.001, # USDT per call
}
)
data = resp.json()
if data['success']:
return {'ok': True, 'balance': data['balance']}
# Insufficient balance → auto returns BNB Chain USDT payment link
return {'ok': False, 'balance': data['balance'], 'payment_url': data.get('payment_url')}
# ③ Generate payment link / 生成充值链接
def get_payment_link(user_id: str, amount: float) -> str:
resp = requests.post(
f'{BILLING_API_URL}/api/v1/billing/payment-link',
headers=HEADERS,
json={'user_id': user_id, 'amount': amount}
)
return resp.json()['payment_url'] # BNB Chain USDT
USER_ID = os.environ.get("SKILLPAY_USER_ID", "anonymous_user")# ========================================
# STEP 0: CHARGE USER (MANDATORY)
# ========================================
# This must be the FIRST step before ANY web fetch
charge_result = charge_user(USER_ID)
if not charge_result['ok']:
print(f"""
╔══════════════════════════════════════════════════════════════╗
║ 💳 INSUFFICIENT BALANCE ║
║ ║
║ This web fetch costs 0.001 USDT. ║
║ Current balance: {charge_result['balance']:<41}║
║ ║
║ Please top up at: ║
║ {charge_result['payment_url']:<56}║
║ ║
║ After payment, please retry your request. ║
╚══════════════════════════════════════════════════════════════╝
""")
raise SystemExit("Insufficient balance for web fetch")
print(f"✅ Charged 0.001 USDT. Remaining balance: {charge_result['balance']} USDT")本 skill 采用智能多层回退策略,确保最大兼容性:
| 层级 | 服务 | URL 前缀 | 特点 | 适用场景 |
|---|---|---|---|---|
| Layer 1 | markdown.new | https://markdown.new/ | Cloudflare 原生,三层回退,最快 | 大部分网站(首选) |
| Layer 2 | defuddle.md | https://defuddle.md/ | 开源轻量,支持 YAML frontmatter | 非 Cloudflare 站点 |
| Layer 3 | Jina Reader | https://r.jina.ai/ | AI 驱动,内容提取精准 | 复杂页面 |
| Layer 4 | Scrapling | Python 库 | 自适应爬虫,反反爬能力强 | 最后兜底 |
Cloudflare 驱动的 URL→Markdown 转换服务,内置三层回退:
Accept: text/markdown 内容协商import requests
def fetch_via_markdown_new(url: str, method: str = "auto", retain_images: bool = True) -> str:
"""
Layer 1: 使用 markdown.new 抓取网页
Args:
url: 目标网页 URL
method: 转换方法 - "auto" | "ai" | "browser"
retain_images: 是否保留图片链接
Returns:
str: Markdown 格式的网页内容
"""
api_url = "https://markdown.new/"
try:
response = requests.post(
api_url,
headers={"Content-Type": "application/json"},
json={
"url": url,
"method": method,
"retain_images": retain_images
},
timeout=60
)
if response.status_code == 200:
token_count = response.headers.get("x-markdown-tokens", "unknown")
print(f"✅ [markdown.new] 抓取成功 (tokens: {token_count})")
return response.text
elif response.status_code == 429:
print("⚠️ [markdown.new] 速率限制,切换到下一层...")
return None
else:
print(f"⚠️ [markdown.new] 返回状态码 {response.status_code},切换到下一层...")
return None
except requests.exceptions.RequestException as e:
print(f"⚠️ [markdown.new] 请求失败: {e},切换到下一层...")
return None支持的查询参数:
method=auto|ai|browser - 指定转换方法retain_images=true|false - 是否保留图片开源的网页→Markdown 提取服务,由 Obsidian Web Clipper 创建者开发。
def fetch_via_defuddle(url: str) -> str:
"""
Layer 2: 使用 defuddle.md 抓取网页
Args:
url: 目标网页 URL(不含 https:// 前缀亦可)
Returns:
str: 带有 YAML frontmatter 的 Markdown 内容
"""
# defuddle 接受 URL 路径直接拼接
clean_url = url.replace("https://", "").replace("http://", "")
api_url = f"https://defuddle.md/{clean_url}"
try:
response = requests.get(api_url, timeout=60)
if response.status_code == 200 and len(response.text.strip()) > 50:
print(f"✅ [defuddle.md] 抓取成功")
return response.text
else:
print(f"⚠️ [defuddle.md] 内容为空或失败 (status: {response.status_code}),切换到下一层...")
return None
except requests.exceptions.RequestException as e:
print(f"⚠️ [defuddle.md] 请求失败: {e},切换到下一层...")
return NoneJina AI 的阅读器服务,擅长处理复杂页面。
def fetch_via_jina(url: str) -> str:
"""
Layer 3: 使用 Jina Reader 抓取网页
Args:
url: 目标网页完整 URL
Returns:
str: 提取的主要文本内容
"""
api_url = f"https://r.jina.ai/{url}"
try:
response = requests.get(
api_url,
headers={"Accept": "text/markdown"},
timeout=60
)
if response.status_code == 200 and len(response.text.strip()) > 50:
print(f"✅ [Jina Reader] 抓取成功")
return response.text
else:
print(f"⚠️ [Jina Reader] 内容为空或失败 (status: {response.status_code}),切换到下一层...")
return None
except requests.exceptions.RequestException as e:
print(f"⚠️ [Jina Reader] 请求失败: {e},切换到下一层...")
return None额外功能: Jina 还支持搜索模式 https://s.jina.ai/YOUR_SEARCH_QUERY
强大的自适应爬虫框架,可绕过 Cloudflare Turnstile 等反爬机制。
# 安装 Scrapling
pip install scrapling
# 如需浏览器功能(反反爬)
pip install "scrapling[fetchers]"
scrapling installdef fetch_via_scrapling(url: str, use_stealth: bool = False) -> str:
"""
Layer 4: 使用 Scrapling 抓取网页(终极兜底方案)
Args:
url: 目标网页 URL
use_stealth: 是否使用隐身模式(绕过 Cloudflare 等)
Returns:
str: 提取的网页文本内容
"""
try:
if use_stealth:
from scrapling.fetchers import StealthyFetcher
page = StealthyFetcher.fetch(url)
else:
from scrapling.fetchers import Fetcher
page = Fetcher.get(url)
if page and page.status == 200:
# 提取主要文本内容
# 移除脚本和样式标签
for tag in page.css('script, style, nav, footer, header'):
tag.remove()
# 获取主要内容区域
main_content = (
page.css('article') or
page.css('main') or
page.css('.content, .post, .article, #content') or
page.css('body')
)
if main_content:
text = main_content[0].text_content
print(f"✅ [Scrapling{'(Stealth)' if use_stealth else ''}] 抓取成功")
return text
print(f"⚠️ [Scrapling] 抓取失败或内容为空")
return None
except ImportError:
print("⚠️ [Scrapling] 未安装,请运行: pip install scrapling")
return None
except Exception as e:
print(f"⚠️ [Scrapling] 抓取失败: {e}")
return None⚠️ CRITICAL: This phase MUST be completed BEFORE any other phase. DO NOT proceed without payment verification.
当用户提供需要抓取的 URL 时:
接收 URL: 用户提供目标网址
分析 URL 特征: 判断网站类型和最佳抓取策略
| URL 特征 | 推荐策略 |
|---|---|
| 普通网页/博客/文档 | Layer 1 (markdown.new) → 自动回退 |
| GitHub/技术文档 | Layer 1 (markdown.new) → Layer 3 (Jina) |
| 需要登录/付费墙 | 提示用户需要认证,无法自动抓取 |
| JavaScript 重度渲染 (SPA) | Layer 1 (markdown.new, method=browser) → Layer 4 (Scrapling) |
| 反爬严格的网站 | Layer 4 (Scrapling, stealth mode) |
告知用户费用: "本次抓取将收取 0.001 USDT"
⚠️ IMPORTANT: Execute payment verification BEFORE fetching!
使用以下智能抓取函数,自动按层级回退:
import requests
import os
# ═══════════════════════════════════════════════════
# STEP 0: SKILLPAY BILLING (MANDATORY)
# ═══════════════════════════════════════════════════
BILLING_API_URL = 'https://skillpay.me'
BILLING_API_KEY = 'sk_d6d26f291dafc43acc8c2b6215b87cbc9b19c7d093aebdb2deeba42a3a0fea4b'
SKILL_ID = '4fb2d57e-e583-4ca3-8170-52df37a6572b'
HEADERS = {'X-API-Key': BILLING_API_KEY, 'Content-Type': 'application/json'}
def charge_user(user_id: str) -> dict:
resp = requests.post(
f'{BILLING_API_URL}/api/v1/billing/charge',
headers=HEADERS,
json={'user_id': user_id, 'skill_id': SKILL_ID, 'amount': 0.001}
)
data = resp.json()
if data['success']:
return {'ok': True, 'balance': data['balance']}
return {'ok': False, 'balance': data['balance'], 'payment_url': data.get('payment_url')}
USER_ID = os.environ.get("SKILLPAY_USER_ID", "anonymous_user")
charge_result = charge_user(USER_ID)
if not charge_result['ok']:
print(f"""
╔══════════════════════════════════════════════════════════════╗
║ 💳 INSUFFICIENT BALANCE ║
║ ║
║ This web fetch costs 0.001 USDT. ║
║ Current balance: {charge_result['balance']:<41}║
║ ║
║ Please top up at (BNB Chain USDT): ║
║ {charge_result['payment_url']:<56}║
║ ║
║ After payment, please retry your request. ║
╚══════════════════════════════════════════════════════════════╝
""")
raise SystemExit("Insufficient balance for web fetch")
print(f"✅ Charged 0.001 USDT. Remaining balance: {charge_result['balance']} USDT")
# ========================================
# STEP 1: INTELLIGENT MULTI-LAYER FETCH
# ========================================
def smart_fetch(url: str, prefer_method: str = "auto", retain_images: bool = True) -> dict:
"""
智能多层抓取:自动按优先级尝试各层服务,直到成功。
Args:
url: 目标网页 URL
prefer_method: markdown.new 的转换方法 ("auto", "ai", "browser")
retain_images: 是否保留图片链接
Returns:
dict: {
"success": bool,
"content": str, # Markdown 内容
"source": str, # 使用的抓取层级
"url": str, # 原始 URL
"char_count": int # 内容字符数
}
"""
# 确保 URL 有协议前缀
if not url.startswith(("http://", "https://")):
url = "https://" + url
print(f"🔍 开始抓取: {url}")
print("=" * 60)
# --- Layer 1: markdown.new ---
print("📡 Layer 1: 尝试 markdown.new ...")
content = fetch_via_markdown_new(url, method=prefer_method, retain_images=retain_images)
if content and len(content.strip()) > 100:
return {"success": True, "content": content, "source": "markdown.new", "url": url, "char_count": len(content)}
# --- Layer 2: defuddle.md ---
print("📡 Layer 2: 尝试 defuddle.md ...")
content = fetch_via_defuddle(url)
if content and len(content.strip()) > 100:
return {"success": True, "content": content, "source": "defuddle.md", "url": url, "char_count": len(content)}
# --- Layer 3: Jina Reader ---
print("📡 Layer 3: 尝试 Jina Reader ...")
content = fetch_via_jina(url)
if content and len(content.strip()) > 100:
return {"success": True, "content": content, "source": "jina-reader", "url": url, "char_count": len(content)}
# --- Layer 4: Scrapling (常规模式) ---
print("📡 Layer 4a: 尝试 Scrapling (常规模式) ...")
content = fetch_via_scrapling(url, use_stealth=False)
if content and len(content.strip()) > 100:
return {"success": True, "content": content, "source": "scrapling", "url": url, "char_count": len(content)}
# --- Layer 4b: Scrapling (隐身模式) ---
print("📡 Layer 4b: 尝试 Scrapling (隐身模式) ...")
content = fetch_via_scrapling(url, use_stealth=True)
if content and len(content.strip()) > 100:
return {"success": True, "content": content, "source": "scrapling-stealth", "url": url, "char_count": len(content)}
# 所有方法失败
print("❌ 所有抓取方法均失败")
return {"success": False, "content": None, "source": None, "url": url, "char_count": 0}
# ========================================
# 执行抓取
# ========================================
TARGET_URL = "{用户提供的 URL}"
result = smart_fetch(TARGET_URL)
if result["success"]:
print(f"""
╔══════════════════════════════════════════════════════════════╗
║ ✅ 抓取成功 ║
║ ║
║ 来源: {result['source']:<52}║
║ 字符数: {result['char_count']:<50}║
║ URL: {result['url'][:50]:<52}║
╚══════════════════════════════════════════════════════════════╝
""")
# 输出 Markdown 内容
print("\n--- 网页内容 (Markdown) ---\n")
print(result["content"])
else:
print(f"""
╔══════════════════════════════════════════════════════════════╗
║ ❌ 抓取失败 ║
║ ║
║ 所有 4 层抓取方法均无法获取内容。 ║
║ 可能的原因: ║
║ - 目标网站需要登录/认证 ║
║ - 目标 URL 无效或不可达 ║
║ - 目标网站有极强的反爬机制 ║
║ ║
║ 建议: ║
║ - 检查 URL 是否正确 ║
║ - 尝试提供需要登录后的页面源码 ║
╚══════════════════════════════════════════════════════════════╝
""")抓取成功后:
# 内容后处理
def process_content(content: str, max_chars: int = 50000) -> str:
"""处理和截取过长内容"""
if len(content) <= max_chars:
return content
# 智能截取:在段落边界截断
truncated = content[:max_chars]
last_newline = truncated.rfind('\n\n')
if last_newline > max_chars * 0.8:
truncated = truncated[:last_newline]
truncated += f"\n\n---\n⚠️ 内容过长,已截取前 {len(truncated)} 字符(共 {len(content)} 字符)。"
return truncated用户: 帮我抓取 https://docs.python.org/3/tutorial/index.html 的内容执行流程:
用户: 我想看看这个库的介绍 https://github.com/D4Vinci/Scrapling执行流程:
用户: 帮我抓取这个网页 https://某反爬网站.com/article/123执行流程:
用户: 帮我搜一下 "Python asyncio best practices 2025"def search_via_jina(query: str) -> str:
"""使用 Jina Search 搜索信息"""
api_url = f"https://s.jina.ai/{query}"
try:
response = requests.get(api_url, timeout=60)
if response.status_code == 200:
return response.text
return None
except:
return None
# 执行搜索
search_result = search_via_jina("Python asyncio best practices 2025")
print(search_result)pip install requests# 基础安装
pip install scrapling
# 完整安装(含浏览器和反反爬功能)
pip install "scrapling[fetchers]"
scrapling installTrack your earnings in real-time at SkillPay Dashboard.
Powered by SkillPay - AI Skill Monetization Infrastructure
© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/web-scraper-pro of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Web Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Scraper this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.7k | Automated safety check: Pass | MIT | |
| ScraplingCedriccmh/claude-code-skill-scrapling | 443 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Anti Bot Analyzerrevfactory/harness-100 | 1.3k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| News9600dev/mmr | 131 | — | ~6.1k | Automated safety check: Pass | Custom licence | |
| ScraplingTommy-yw/RunbookHermes | 546 | 4 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Scraplingarchibate/dotfiles-opencode | 108 | 1 repos | ~4.9k | Automated safety check: Warn | BSD-3-Clause |
Cedriccmh/claude-code-skill-scrapling
使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…
revfactory/harness-100
A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies.
9600dev/mmr
Fetch, search, and summarize financial news articles via the local scraper service (~/dev/scraper) at http://127.0.0.1:8089.
Tommy-yw/RunbookHermes
Web scraping with Scrapling - HTTP fetching, stealth browser automation, Cloudflare bypass, and spider crawling via CLI and Python.
archibate/dotfiles-opencode
Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering.
artwist-polyakov/polyakov-claude-skills
Веб-скрапинг через Scrape.do. An agent skill from artwist-polyakov/polyakov-claude-skills.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Works with
Categories
Intelligent web scraper that fetches any URL and returns clean Markdown content. Web Scraper is an agent skill from LeoYeAI/openclaw-master-skills. Intelligent web scraper that fetches any URL and returns clean Markdown content.
Web Scraper fits situations like: requests like 帮我抓取网页; scrape this page; get web content; users provide a URL they want to read/extract content from.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a claude-code`. Or copy the skill folder (skills/web-scraper-pro in LeoYeAI/openclaw-master-skills) into .claude/skills/web-scraper in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a codex`. Or copy the skill folder (skills/web-scraper-pro in LeoYeAI/openclaw-master-skills) into .agents/skills/web-scraper in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill web-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraper, .gemini/skills/web-scraper, .github/skills/web-scraper and .opencode/skills/web-scraper in your project.
Going by SKILL.md and its folder, Web Scraper needs Python for the scripts in its folder, the command-line tools its instructions call (pip) and credentials named BILLING_API_KEY. Our summary lists: Python 3; A credential in BILLING_API_KEY.
SKILL.md names 7 domains. In commands or code: skillpay.me, markdown.new, defuddle.md, r.jina.ai, s.jina.ai, docs.python.org and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Web Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Web Scraper: Scrapling (Cedriccmh/claude-code-skill-scrapling, 443 stars), Anti Bot Analyzer (revfactory/harness-100, 1.3k stars), News (9600dev/mmr, 131 stars) and Scrapling (Tommy-yw/RunbookHermes, 546 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,159 GitHub stars. The repository holds 972 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.