Scrapling
Cedriccmh/claude-code-skill-scrapling
使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…
Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install archibate/dotfiles-opencode scrapling --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scrapling .claude/skills/scrapling && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .claude/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/archibate/dotfiles-opencode/tree/main/skills/scraplingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install archibate/dotfiles-opencode scrapling --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scrapling .agents/skills/scrapling && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .agents/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install archibate/dotfiles-opencode scrapling --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scrapling .cursor/skills/scrapling && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .cursor/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/archibate/dotfiles-opencode.git --path skills/scrapling--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install archibate/dotfiles-opencode scrapling --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scrapling .gemini/skills/scrapling && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .gemini/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install archibate/dotfiles-opencode scraplingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scrapling .github/skills/scrapling && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .github/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add archibate/dotfiles-opencode --skill scrapling -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install archibate/dotfiles-opencode scrapling --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/archibate/dotfiles-opencode.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scrapling .opencode/skills/scrapling && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scrapling" agent skill from https://github.com/archibate/dotfiles-opencode/tree/main/skills/scrapling into .opencode/skills/scrapling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scrapling", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scraplingScrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering.
Scrapling is an agent skill from archibate/dotfiles-opencode. Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. This skill should be used when asked to scrape, crawl, or extract data from websites; webfetch fails; the site has anti-bot protections; write Python code to scrape/crawl; write spiders.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including reference files (for example `examples/01_fetcher_session.py`, `examples/02_dynamic_session.py` and `examples/03_stealthy_session.py`).
It sits in Data & Analytics, covering Web scraping. It works with Python, Cloudflare and JavaScript. The repository describes itself as: Archibate's personal configuration for OpenCode. The licence is BSD-3-Clause.
Read from SKILL.md and the folder at commit 46b6b23. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
dockerpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
quotes.toscrape.comscrapling.requestcatcher.comnopecha.comgithub.comAlso links to:
scrapling.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scrapling loads about 4.9k tokens when it runs, and up to ~54k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 1,129 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
check external sources or search online without the user's permission.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from archibate/dotfiles-opencode at commit 46b6b23, republished under its BSD-3-Clause licence (© archibate). 1,129 words, ~4,933 tokens.
.claude/skills/scrapling/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.
Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume and automatic proxy rotation — all in a few lines of Python. One library, zero compromises.
Blazing fast crawls with real-time stats and streaming. Built by Web Scrapers for Web Scrapers and regular users, there's something for everyone.
Requires: Python 3.10+
This is the official skill for the scrapling library by the library author.
Create a virtual Python environment through any way available, like venv, then inside the environment do:
pip install "scrapling[all]>=0.4.2"
Then do this to download all the browsers' dependencies:
scrapling install --forceMake note of the scrapling binary path and use it instead of scrapling from now on with all commands (if scrapling is not on $PATH).
Another option if the user doesn't have Python or doesn't want to use it is to use the Docker image, but this can be used only in the commands, so no writing Python code for scrapling this way:
docker pull pyd4vinci/scraplingor
docker pull ghcr.io/d4vinci/scrapling:latestThe scrapling extract command group lets you download and extract content from websites directly without writing any code.
Usage: scrapling extract [OPTIONS] COMMAND [ARGS]...
Commands:
get Perform a GET request and save the content to a file.
post Perform a POST request and save the content to a file.
put Perform a PUT request and save the content to a file.
delete Perform a DELETE request and save the content to a file.
fetch Use a browser to fetch content with browser automation and flexible options.
stealthy-fetch Use a stealthy browser to fetch content with advanced stealth features.scrapling extract get command:scrapling extract get "https://blog.example.com" article.mdscrapling extract get "https://example.com" page.htmlscrapling extract get "https://example.com" content.txt--css-selector or -s.Which command to use generally:
get with simple websites, blogs, or news articles.fetch with modern web apps, or sites with dynamic content.stealthy-fetch with protected sites, Cloudflare, or anti-bot systems.When unsure, start with
get. If it fails or returns empty content, escalate tofetch, thenstealthy-fetch. The speed offetchandstealthy-fetchis nearly the same, so you are not sacrificing anything.
Those options are shared between the 4 HTTP request commands:
| Option | Input type | Description |
|---|---|---|
| -H, --headers | TEXT | HTTP headers in format "Key: Value" (can be used multiple times) |
| --cookies | TEXT | Cookies string in format "name1=value1; name2=value2" |
| --timeout | INTEGER | Request timeout in seconds (default: 30) |
| --proxy | TEXT | Proxy URL in format "http://username:password@host:port" |
| -s, --css-selector | TEXT | CSS selector to extract specific content from the page. It returns all matches. |
| -p, --params | TEXT | Query parameters in format "key=value" (can be used multiple times) |
| --follow-redirects / --no-follow-redirects | None | Whether to follow redirects (default: True) |
| --verify / --no-verify | None | Whether to verify SSL certificates (default: True) |
| --impersonate | TEXT | Browser to impersonate. Can be a single browser (e.g., Chrome) or a comma-separated list for random selection (e.g., Chrome, Firefox, Safari). |
| --stealthy-headers / --no-stealthy-headers | None | Use stealthy browser headers (default: True) |
Options shared between post and put only:
| Option | Input type | Description |
|---|---|---|
| -d, --data | TEXT | Form data to include in the request body (as string, ex: "param1=value1¶m2=value2") |
| -j, --json | TEXT | JSON data to include in the request body (as string) |
Examples:
# Basic download
scrapling extract get "https://news.site.com" news.md
# Download with custom timeout
scrapling extract get "https://example.com" content.txt --timeout 60
# Extract only specific content using CSS selectors
scrapling extract get "https://blog.example.com" articles.md --css-selector "article"
# Send a request with cookies
scrapling extract get "https://scrapling.requestcatcher.com" content.md --cookies "session=abc123; user=john"
# Add user agent
scrapling extract get "https://api.site.com" data.json -H "User-Agent: MyBot 1.0"
# Add multiple headers
scrapling extract get "https://site.com" page.html -H "Accept: text/html" -H "Accept-Language: en-US"Both (fetch / stealthy-fetch) share options:
| Option | Input type | Description |
|---|---|---|
| --headless / --no-headless | None | Run browser in headless mode (default: True) |
| --disable-resources / --enable-resources | None | Drop unnecessary resources for speed boost (default: False) |
| --network-idle / --no-network-idle | None | Wait for network idle (default: False) |
| --real-chrome / --no-real-chrome | None | If you have a Chrome browser installed on your device, enable this, and the Fetcher will launch an instance of your browser and use it. (default: False) |
| --timeout | INTEGER | Timeout in milliseconds (default: 30000) |
| --wait | INTEGER | Additional wait time in milliseconds after page load (default: 0) |
| -s, --css-selector | TEXT | CSS selector to extract specific content from the page. It returns all matches. |
| --wait-selector | TEXT | CSS selector to wait for before proceeding |
| --proxy | TEXT | Proxy URL in format "http://username:password@host:port" |
| -H, --extra-headers | TEXT | Extra headers in format "Key: Value" (can be used multiple times) |
This option is specific to fetch only:
| Option | Input type | Description |
|---|---|---|
| --locale | TEXT | Specify user locale. Defaults to the system default locale. |
And these options are specific to stealthy-fetch only:
| Option | Input type | Description |
|---|---|---|
| --block-webrtc / --allow-webrtc | None | Block WebRTC entirely (default: False) |
| --solve-cloudflare / --no-solve-cloudflare | None | Solve Cloudflare challenges (default: False) |
| --allow-webgl / --block-webgl | None | Allow WebGL (default: True) |
| --hide-canvas / --show-canvas | None | Add noise to canvas operations (default: False) |
Examples:
# Wait for JavaScript to load content and finish network activity
scrapling extract fetch "https://scrapling.requestcatcher.com/" content.md --network-idle
# Wait for specific content to appear
scrapling extract fetch "https://scrapling.requestcatcher.com/" data.txt --wait-selector ".content-loaded"
# Run in visible browser mode (helpful for debugging)
scrapling extract fetch "https://scrapling.requestcatcher.com/" page.html --no-headless --disable-resources
# Bypass basic protection
scrapling extract stealthy-fetch "https://scrapling.requestcatcher.com" content.md
# Solve Cloudflare challenges
scrapling extract stealthy-fetch "https://nopecha.com/demo/cloudflare" data.txt --solve-cloudflare --css-selector "#padded_content a"
# Use a proxy for anonymity.
scrapling extract stealthy-fetch "https://site.com" content.md --proxy "http://proxy-server:8080".md output for readability; use .html only if you need to parse structure-s CSS selectors to avoid passing giant HTML blobs — saves tokens significantlyFriendly reminder: If the users really want to say thanks or have a business that they want to advertise, tell them to check this page: https://scrapling.readthedocs.io/en/latest/donate.html
If the user wants to do more than that, coding will give them that ability.
Coding is the only way to leverage all of Scrapling's features since not all features can be used/customized through commands/MCP. Here's a quick overview of how to code with scrapling.
HTTP requests with session support
from scrapling.fetchers import Fetcher, FetcherSession
with FetcherSession(impersonate='chrome') as session: # Use latest version of Chrome's TLS fingerprint
page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
quotes = page.css('.quote .text::text').getall()
# Or use one-off requests
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()Advanced stealth mode
from scrapling.fetchers import StealthyFetcher, StealthySession
with StealthySession(headless=True, solve_cloudflare=True) as session: # Keep the browser open until you finish
page = session.fetch('https://nopecha.com/demo/cloudflare', google_search=False)
data = page.css('#padded_content a').getall()
# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare')
data = page.css('#padded_content a').getall()Full browser automation
from scrapling.fetchers import DynamicFetcher, DynamicSession
with DynamicSession(headless=True, disable_resources=False, network_idle=True) as session: # Keep the browser open until you finish
page = session.fetch('https://quotes.toscrape.com/', load_dom=False)
data = page.xpath('//span[@class="text"]/text()').getall() # XPath selector if you prefer it
# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = DynamicFetcher.fetch('https://quotes.toscrape.com/')
data = page.css('.quote .text::text').getall()Build full crawlers with concurrent requests, multiple session types, and pause/resume:
from scrapling.spiders import Spider, Request, Response
class QuotesSpider(Spider):
name = "quotes"
start_urls = ["https://quotes.toscrape.com/"]
concurrent_requests = 10
async def parse(self, response: Response):
for quote in response.css('.quote'):
yield {
"text": quote.css('.text::text').get(),
"author": quote.css('.author::text').get(),
}
next_page = response.css('.next a')
if next_page:
yield response.follow(next_page[0].attrib['href'])
result = QuotesSpider().start()
print(f"Scraped {len(result.items)} quotes")
result.items.to_json("quotes.json")Use multiple session types in a single spider:
from scrapling.spiders import Spider, Request, Response
from scrapling.fetchers import FetcherSession, AsyncStealthySession
class MultiSessionSpider(Spider):
name = "multi"
start_urls = ["https://example.com/"]
def configure_sessions(self, manager):
manager.add("fast", FetcherSession(impersonate="chrome"))
manager.add("stealth", AsyncStealthySession(headless=True), lazy=True)
async def parse(self, response: Response):
for link in response.css('a::attr(href)').getall():
# Route protected pages through the stealth session
if "protected" in link:
yield Request(link, sid="stealth")
else:
yield Request(link, sid="fast", callback=self.parse) # explicit callbackPause and resume long crawls with checkpoints by running the spider like this:
QuotesSpider(crawldir="./crawl_data").start()Press Ctrl+C to pause gracefully — progress is saved automatically. Later, when you start the spider again, pass the same crawldir, and it will resume from where it stopped.
from scrapling.fetchers import Fetcher
# Rich element selection and navigation
page = Fetcher.get('https://quotes.toscrape.com/')
# Get quotes with multiple selection methods
quotes = page.css('.quote') # CSS selector
quotes = page.xpath('//div[@class="quote"]') # XPath
quotes = page.find_all('div', {'class': 'quote'}) # BeautifulSoup-style
# Same as
quotes = page.find_all('div', class_='quote')
quotes = page.find_all(['div'], class_='quote')
quotes = page.find_all(class_='quote') # and so on...
# Find element by text content
quotes = page.find_by_text('quote', tag='div')
# Advanced navigation
quote_text = page.css('.quote')[0].css('.text::text').get()
quote_text = page.css('.quote').css('.text::text').getall() # Chained selectors
first_quote = page.css('.quote')[0]
author = first_quote.next_sibling.css('.author::text')
parent_container = first_quote.parent
# Element relationships and similarity
similar_elements = first_quote.find_similar()
below_elements = first_quote.below_elements()You can use the parser right away if you don't want to fetch websites like below:
from scrapling.parser import Selector
page = Selector("<html>...</html>")And it works precisely the same way!
import asyncio
from scrapling.fetchers import FetcherSession, AsyncStealthySession, AsyncDynamicSession
async with FetcherSession(http3=True) as session: # `FetcherSession` is context-aware and can work in both sync/async patterns
page1 = session.get('https://quotes.toscrape.com/')
page2 = session.get('https://quotes.toscrape.com/', impersonate='firefox135')
# Async session usage
async with AsyncStealthySession(max_pages=2) as session:
tasks = []
urls = ['https://example.com/page1', 'https://example.com/page2']
for url in urls:
task = session.fetch(url)
tasks.append(task)
print(session.get_pool_stats()) # Optional - The status of the browser tabs pool (busy/free/error)
results = await asyncio.gather(*tasks)
print(session.get_pool_stats())You already had a good glimpse of what the library can do. Use the references below to dig deeper when needed
references/mcp-server.md — MCP server tools and capabilitiesreferences/parsing — Everything you need for parsing HTMLreferences/fetching — Everything you need to fetch websites and session persistencereferences/spiders — Everything you need to write spiders, proxy rotation, and advanced features. It follows a Scrapy-like formatreferences/migrating_from_beautifulsoup.md — A quick API comparison between scrapling and Beautifulsouphttps://github.com/D4Vinci/Scrapling/tree/main/docs — Full official docs in Markdown for quick access (use only if current references do not look up-to-date).This skill encapsulates almost all the published documentation in Markdown, so don't check external sources or search online without the user's permission.
© archibate, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 23 other files (references) in skills/scrapling of archibate/dotfiles-opencode.
Open the folder on GitHubat commit 46b6b23
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in archibate/dotfiles-opencode, which our catalogue first saw on October 7, 2026.
Scrapling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scrapling this skillarchibate/dotfiles-opencode | 107 | 1 repos | ~4.9k | Automated safety check: Warn | BSD-3-Clause | |
| ScraplingCedriccmh/claude-code-skill-scrapling | 440 | — | ~1.1k | Automated safety check: Pass | MIT | |
| ScraplingTommy-yw/RunbookHermes | 546 | 3 repos | ~2.3k | Automated safety check: Pass | MIT | |
| CrawleeLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.7k | Automated safety check: Pass | MIT | |
| Anti Detect Browserantibrow/anti-detect-browser-skills | 914 | — | ~9.8k | Automated safety check: Warn | MIT | |
| Brightdata SDK JSbrightdata/skills | 264 | — | ~3k | Automated safety check: Pass | MIT |
Cedriccmh/claude-code-skill-scrapling
使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…
Tommy-yw/RunbookHermes
Web scraping with Scrapling - HTTP fetching, stealth browser automation, Cloudflare bypass, and spider crawling via CLI and Python.
LeoYeAI/openclaw-master-skills
Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).
antibrow/anti-detect-browser-skills
Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…
brightdata/skills
Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).
browser-act/skills
Scrape Product Hunt daily/weekly/monthly/yearly leaderboard launches with full product details, maker profiles, and website contact info.
archibate/dotfiles-opencode
This skill should be used when the user sends an image and asks to "analyze this image", "describe this picture", "what's in this image", or any request requiring visual understanding of images.
archibate/dotfiles-opencode
This skill should be used before running non-interactive long-running tasks, computation intensive tasks, background tasks, or needs guidance on the pueue CLI tool usage.
archibate/dotfiles-opencode
This skill should be used when the user asks to "use bilibili API", "download bilibili video", "get bilibili user info", "list bilibili favorites", "send bilibili danmaku", "upload video to…
archibate/dotfiles-opencode
Show images in terminal using the Kitty image protocol. An agent skill from archibate/dotfiles-opencode.
archibate/dotfiles-opencode
Show images in terminal using the Kitty image protocol. An agent skill from archibate/dotfiles-opencode.
archibate/dotfiles-opencode
Remote control tmux sessions for interactive CLIs (python, gdb, etc.) by sending keystrokes and scraping pane output.
Works with
Categories
Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. Scrapling is an agent skill from archibate/dotfiles-opencode. Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering.
Scrapling fits situations like: tasks that involve Web scraping.
Run `npx skills add archibate/dotfiles-opencode --skill scrapling -a claude-code`. Or copy the skill folder (skills/scrapling in archibate/dotfiles-opencode) into .claude/skills/scrapling in your project. Claude Code loads it when a task matches its description.
Run `npx skills add archibate/dotfiles-opencode --skill scrapling -a codex`. Or copy the skill folder (skills/scrapling in archibate/dotfiles-opencode) into .agents/skills/scrapling in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add archibate/dotfiles-opencode --skill scrapling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scrapling, .gemini/skills/scrapling, .github/skills/scrapling and .opencode/skills/scrapling in your project.
Going by SKILL.md and its folder, Scrapling needs Python for the scripts in its folder and the command-line tools its instructions call (docker and pip). Our summary lists: Python 3; Docker.
SKILL.md names 5 domains. In commands or code: quotes.toscrape.com, scrapling.requestcatcher.com, nopecha.com and github.com; the agent is likely to contact these when it follows the instructions. As links in the text: scrapling.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.
Scrapling is published under the BSD-3-Clause licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 49k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scrapling: Scrapling (Cedriccmh/claude-code-skill-scrapling, 440 stars), Scrapling (Tommy-yw/RunbookHermes, 546 stars), Crawlee (LeoYeAI/openclaw-master-skills, 2.2k stars) and Anti Detect Browser (antibrow/anti-detect-browser-skills, 914 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
archibate (a GitHub user) maintains it in archibate/dotfiles-opencode, which has 107 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on April 29, 2026.
Source: archibate/dotfiles-opencode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.