Scrape
davila7/claude-code-templates
Scrape any webpage as clean markdown via Bright Data Web Unlocker API.
Build production-ready web scrapers for any website using Bright Data infrastructure.
$ npx skills add brightdata/skills --skill scraper-builder -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install brightdata/skills scraper-builder --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scraper-builder .claude/skills/scraper-builder && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .claude/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/brightdata/skills/tree/main/skills/scraper-builderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add brightdata/skills --skill scraper-builder -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install brightdata/skills scraper-builder --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scraper-builder .agents/skills/scraper-builder && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .agents/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brightdata/skills --skill scraper-builder -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install brightdata/skills scraper-builder --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scraper-builder .cursor/skills/scraper-builder && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .cursor/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/brightdata/skills.git --path skills/scraper-builder--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add brightdata/skills --skill scraper-builder -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install brightdata/skills scraper-builder --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scraper-builder .gemini/skills/scraper-builder && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .gemini/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install brightdata/skills scraper-builderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add brightdata/skills --skill scraper-builder -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scraper-builder .github/skills/scraper-builder && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .github/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brightdata/skills --skill scraper-builder -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install brightdata/skills scraper-builder --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scraper-builder .opencode/skills/scraper-builder && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scraper-builder" agent skill from https://github.com/brightdata/skills/tree/main/skills/scraper-builder into .opencode/skills/scraper-builder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scraper-builder", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scraper-builderBuild production-ready web scrapers for any website using Bright Data infrastructure.
Scraper Builder is an agent skill from brightdata/skills. Build production-ready web scrapers for any website using Bright Data infrastructure. Guides you through site analysis, API selection, selector extraction, pagination handling, and complete scraper implementation. Use this skill whenever the user wants to build a scraper, create a crawler, extract data from a website, scrape product pages, handle pagination, build a data pipeline from a web source, or automate data collection from any site — even if they don't explicitly say 'scraper'. Triggers on phrases like…
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/evals.json`, `references/concurrency-guide.md` and `references/pagination-patterns.md`).
It sits in Data & Analytics, covering Web scraping. It works with Bright Data. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bashcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.brightdata.comamazon.comtarget-site.comdocs.brightdata.combrightdata.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BRIGHTDATA_API_KEYAPI_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scraper Builder loads about 7.2k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 2,200 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 2,200 words, ~7,195 tokens.
.claude/skills/scraper-builder/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You are building a production-ready web scraper for the user. Your job is to guide them from "I want data from site X" to a working, robust scraper that handles real-world challenges like pagination, dynamic content, anti-bot protection, and data parsing.
After building the scraper, always run it on a small sample (1-3 pages) and show the extracted data to the user before scaling up. If the output is empty, malformed, or missing fields, iterate — fix selectors, switch APIs, or adjust the parsing logic. A scraper that doesn't produce clean data is not done.
Take your time with the reconnaissance phase. Spending 2 minutes analyzing the HTML upfront prevents hours of debugging later. Quality is more important than speed here.
This skill orchestrates Bright Data's four APIs to build scrapers intelligently. Rather than writing fragile custom scraping code, you analyze the target site first, then pick the most reliable and cost-effective extraction method. The decision tree is:
The skill produces complete, runnable code — not pseudocode or outlines.
Before writing any code, you need to understand what the user wants and what the site looks like. Ask these questions (skip any the user already answered):
Don't over-interview. If the user says "build a scraper for Amazon product pages", you already know: site=Amazon, data=product details, scope=product pages. Jump ahead.
Before doing any custom work, check if Bright Data already has a scraper for this domain. This is the fastest, cheapest, and most reliable path.
Read references/supported-domains.md for the curated list of common pre-built scrapers. But the curated list may not be complete — Bright Data supports 100+ domains and adds new scrapers regularly. If you don't see the target domain in the curated list, query the live Dataset List API to check:
curl -H "Authorization: Bearer $BRIGHTDATA_API_KEY" \
https://api.brightdata.com/datasets/listThis returns every available scraper with its dataset_id and name. Search the results for the target domain. You can also browse the full documentation index at https://docs.brightdata.com/llms.txt to discover scraper-specific docs and supported parameters.
Use the Web Scraper API or Python SDK platform-specific scrapers. This gives you structured JSON with no parsing code needed.
Python SDK approach (preferred):
from brightdata import BrightDataClient
async with BrightDataClient() as client:
result = await client.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
if result.success:
print(result.data) # Structured product dataREST API approach (shell/curl):
bash scripts/datasets.sh amazon_product "https://www.amazon.com/dp/B09V3KXJPB"For bulk scraping with pre-built scrapers, use the async trigger/poll/fetch pattern:
async with BrightDataClient() as client:
# Trigger without waiting
job = await client.scrape.amazon.products_trigger(url=url)
# Poll until ready
await job.wait(timeout=180, poll_interval=10, verbose=True)
# Fetch results
data = await job.fetch()Skip to Phase 5 (pagination/orchestration) if the user needs multi-page scraping with a pre-built scraper.
Continue to Phase 3 — you need to analyze the site and build a custom scraper.
This is the critical step that separates reliable scrapers from brittle ones. You need to understand the site's structure before writing extraction code.
Use Web Unlocker to get the raw HTML. This tells you whether the content is server-rendered or client-rendered, and gives you the actual DOM to analyze.
import requests
import os
API_KEY = os.environ["BRIGHTDATA_API_KEY"]
ZONE = os.environ["BRIGHTDATA_UNLOCKER_ZONE"]
response = requests.post(
"https://api.brightdata.com/request",
headers={"Authorization": f"Bearer {API_KEY}"},
json={
"zone": ZONE,
"url": "https://target-site.com/page",
"format": "raw"
}
)
html = response.textOr use the scrape skill's shell script:
bash skills/scrape/scripts/scrape.sh "https://target-site.com/page"Read references/site-analysis-guide.md for the detailed analysis playbook.
Look at the fetched HTML and determine:
Is the content in the HTML? If the data you need is present in the raw HTML, Web Unlocker is sufficient. If the HTML is mostly empty shells with JS framework markers (<div id="root"></div>, <div id="__next"></div>, ng-app), the content is client-rendered and you need Browser API.
Identify reliable selectors. Find the CSS selectors or data attributes that target the data fields. Prefer selectors in this order (most reliable → least):
data-* attributes (e.g., [data-testid="product-price"]) — survive redesigns.product-card .price)id attributes — unique but may changediv > span:nth-child(2)) — fragile, avoidIdentify the data pattern. Is it:
Check for hidden APIs. Many modern sites load data via XHR/fetch calls to internal APIs. If you see structured JSON endpoints in the page source or network activity, hitting those directly through Web Unlocker is often cleaner than parsing HTML.
Based on your analysis:
| Finding | Approach |
|---|---|
| Content in HTML, no interaction needed | Web Unlocker — fetch HTML, parse with BeautifulSoup/Cheerio |
| Content loaded via JSON API | Web Unlocker — hit the API endpoint directly |
| Content requires JS rendering | Browser API — render then extract |
| Content needs click/scroll/interaction | Browser API — automate the interaction |
| Infinite scroll pagination | Browser API — scroll and collect |
| Standard URL-based pagination | Web Unlocker — iterate page URLs |
| CAPTCHA-heavy site | Browser API — auto-solves CAPTCHAs |
Now write the actual extraction code. The approach depends on Phase 3's decision.
Best for static sites or sites with server-rendered HTML. This is the cheapest and fastest approach.
import requests
import os
from bs4 import BeautifulSoup
API_KEY = os.environ["BRIGHTDATA_API_KEY"]
ZONE = os.environ["BRIGHTDATA_UNLOCKER_ZONE"]
def fetch_page(url: str) -> str:
"""Fetch a page through Bright Data Web Unlocker."""
response = requests.post(
"https://api.brightdata.com/request",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"zone": ZONE, "url": url, "format": "raw"}
)
response.raise_for_status()
return response.text
def parse_products(html: str) -> list[dict]:
"""Extract product data from HTML. Customize selectors per site."""
soup = BeautifulSoup(html, "html.parser")
products = []
for card in soup.select(".product-card"): # Adjust selector
product = {
"name": card.select_one(".product-title").get_text(strip=True),
"price": card.select_one(".product-price").get_text(strip=True),
"url": card.select_one("a")["href"],
# Add more fields as needed
}
products.append(product)
return products
# Usage
html = fetch_page("https://example.com/products")
products = parse_products(html)Key patterns for robust parsing:
.get_text(strip=True) to clean whitespace.get("href", "") instead of ["href"] to avoid KeyError on missing attributesWhen you discover the site loads data from a JSON API endpoint, hit it directly. This is the cleanest approach — no HTML parsing needed.
import requests
import json
import os
API_KEY = os.environ["BRIGHTDATA_API_KEY"]
ZONE = os.environ["BRIGHTDATA_UNLOCKER_ZONE"]
def fetch_api(api_url: str) -> dict:
"""Fetch a JSON API endpoint through Web Unlocker."""
response = requests.post(
"https://api.brightdata.com/request",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"zone": ZONE, "url": api_url, "format": "raw"}
)
return json.loads(response.text)
# Example: site with internal API
data = fetch_api("https://example.com/api/products?page=1&limit=50")
products = data["results"] # Already structured!Use when the site requires JavaScript rendering, interaction (clicks, scrolls, form fills), or has aggressive anti-bot measures.
import asyncio
from playwright.async_api import async_playwright
AUTH = os.environ.get("BROWSER_AUTH", "brd-customer-CUSTOMER_ID-zone-ZONE_NAME:PASSWORD")
async def scrape_with_browser(url: str) -> str:
"""Scrape a page using Bright Data Browser API."""
async with async_playwright() as p:
browser = await p.chromium.connect_over_cdp(
f"wss://{AUTH}@brd.superproxy.io:9222"
)
page = await browser.new_page()
page.set_default_navigation_timeout(120_000) # 2 minutes — required
# Block unnecessary resources to reduce bandwidth costs
await page.route("**/*.{png,jpg,jpeg,gif,svg,css,woff,woff2}",
lambda route: route.abort())
await page.goto(url, wait_until="domcontentloaded")
# Wait for the content you need to appear
await page.wait_for_selector(".product-card", timeout=30_000)
# Extract data using page.evaluate for performance
products = await page.evaluate("""
() => Array.from(document.querySelectorAll('.product-card')).map(card => ({
name: card.querySelector('.product-title')?.textContent?.trim(),
price: card.querySelector('.product-price')?.textContent?.trim(),
url: card.querySelector('a')?.href,
}))
""")
await browser.close()
return productsBrowser API rules you must follow:
set_default_navigation_timeout(120_000))page.goto() per session — for a new URL, create a new browser connectionwait_until="domcontentloaded" not networkidle (SPAs never reach networkidle)page.evaluate() for bulk extraction — it's faster than individual selector callsFor sites that load more content when you scroll down.
async def scrape_infinite_scroll(url: str, max_items: int = 100) -> list:
"""Scrape a page with infinite scroll."""
async with async_playwright() as p:
browser = await p.chromium.connect_over_cdp(
f"wss://{AUTH}@brd.superproxy.io:9222"
)
page = await browser.new_page()
page.set_default_navigation_timeout(120_000)
await page.route("**/*.{png,jpg,jpeg,gif,svg,woff,woff2}",
lambda route: route.abort())
await page.goto(url, wait_until="domcontentloaded")
all_items = []
previous_count = 0
while len(all_items) < max_items:
# Scroll to bottom
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(2000) # Wait for content to load
# Extract all currently visible items
items = await page.evaluate("""
() => Array.from(document.querySelectorAll('.item-selector')).map(el => ({
// ... extract fields
}))
""")
all_items = items
if len(all_items) == previous_count:
break # No new content loaded — we've reached the end
previous_count = len(all_items)
await browser.close()
return all_items[:max_items]Most scraping tasks involve multiple pages. The approach depends on the pagination type.
Read references/pagination-patterns.md for detailed pagination strategies.
Pages are accessed via URL parameters like ?page=2 or ?offset=20.
import time
def scrape_all_pages(base_url: str, max_pages: int = 50) -> list[dict]:
"""Scrape all pages of a paginated listing."""
all_items = []
for page_num in range(1, max_pages + 1):
url = f"{base_url}?page={page_num}"
html = fetch_page(url)
items = parse_products(html)
if not items:
break # No more results
all_items.extend(items)
print(f"Page {page_num}: {len(items)} items (total: {len(all_items)})")
time.sleep(1) # Be respectful — don't hammer the site
return all_itemsFollow "next" links found in the HTML.
def scrape_with_next_links(start_url: str) -> list[dict]:
"""Follow next-page links to scrape all pages."""
all_items = []
url = start_url
while url:
html = fetch_page(url)
items = parse_products(html)
all_items.extend(items)
# Find next page link
soup = BeautifulSoup(html, "html.parser")
next_link = soup.select_one("a.next-page, a[rel='next'], .pagination .next a")
url = next_link["href"] if next_link else None
# Handle relative URLs
if url and not url.startswith("http"):
from urllib.parse import urljoin
url = urljoin(start_url, url)
time.sleep(1)
return all_itemsWhen you have many URLs (50+), always use concurrent requests with a semaphore — never fetch them one-by-one in a sequential loop. Read references/concurrency-guide.md for the full concurrency playbook including per-site tuning, multi-site parallelism, and retry strategies.
import asyncio
import aiohttp
CONCURRENCY = 20 # Start here, tune per site — see concurrency guide
async def scrape_pages_concurrent(urls: list[str]) -> list[dict]:
"""Scrape multiple pages with controlled concurrency."""
sem = asyncio.Semaphore(CONCURRENCY)
async def fetch_one(session, url):
async with sem:
async with session.post(
"https://api.brightdata.com/request",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"zone": ZONE, "url": url, "format": "raw"},
timeout=aiohttp.ClientTimeout(total=60),
) as resp:
return {"url": url, "html": await resp.text()}
async with aiohttp.ClientSession() as session:
tasks = [fetch_one(session, url) for url in urls]
results = await asyncio.gather(*tasks, return_exceptions=True)
all_items = []
for r in results:
if not isinstance(r, Exception):
all_items.extend(parse_products(r["html"]))
return all_items
# Generate all page URLs
urls = [f"https://example.com/products?page={i}" for i in range(1, 51)]
items = asyncio.run(scrape_pages_concurrent(urls))Some APIs use cursor tokens instead of page numbers.
def scrape_with_cursor(api_base: str) -> list[dict]:
"""Handle cursor-based API pagination."""
all_items = []
cursor = None
while True:
url = f"{api_base}?limit=100"
if cursor:
url += f"&cursor={cursor}"
data = fetch_api(url)
all_items.extend(data["results"])
cursor = data.get("next_cursor")
if not cursor:
break
return all_itemsNow put it all together into a clean, runnable script. Every scraper you build should have:
Important: If the user has more than ~50 URLs to scrape, the scraper must use concurrent requests — not a sequential loop. See references/concurrency-guide.md for the complete concurrent scraper template and tuning guidelines.
#!/usr/bin/env python3
"""
Scraper for [SITE NAME] - [DESCRIPTION]
Built with Bright Data [API NAME]
Usage:
export BRIGHTDATA_API_KEY="your-api-key"
export BRIGHTDATA_UNLOCKER_ZONE="your-zone-name"
python scraper.py
"""
import json
import os
import sys
import time
import logging
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
# --- Configuration ---
API_KEY = os.environ["BRIGHTDATA_API_KEY"]
ZONE = os.environ["BRIGHTDATA_UNLOCKER_ZONE"]
TARGET_URL = "https://example.com/products"
OUTPUT_FILE = "results.json"
MAX_PAGES = 50
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(message)s")
log = logging.getLogger(__name__)
# --- Fetcher ---
def fetch_page(url: str, retries: int = 3) -> str:
for attempt in range(retries):
try:
response = requests.post(
"https://api.brightdata.com/request",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"zone": ZONE, "url": url, "format": "raw"},
timeout=60,
)
response.raise_for_status()
return response.text
except requests.RequestException as e:
log.warning(f"Attempt {attempt + 1} failed for {url}: {e}")
if attempt < retries - 1:
time.sleep(2 ** attempt)
raise RuntimeError(f"Failed to fetch {url} after {retries} attempts")
# --- Parser ---
def parse_items(html: str) -> list[dict]:
soup = BeautifulSoup(html, "html.parser")
items = []
for card in soup.select("ITEM_SELECTOR"):
try:
item = {
"name": card.select_one("NAME_SELECTOR").get_text(strip=True),
"price": card.select_one("PRICE_SELECTOR").get_text(strip=True),
"url": card.select_one("a").get("href", ""),
}
items.append(item)
except (AttributeError, TypeError) as e:
log.warning(f"Failed to parse item: {e}")
continue
return items
# --- Paginator ---
def scrape_all(base_url: str) -> list[dict]:
all_items = []
for page in range(1, MAX_PAGES + 1):
url = f"{base_url}?page={page}"
log.info(f"Scraping page {page}...")
html = fetch_page(url)
items = parse_items(html)
if not items:
log.info(f"No items on page {page} — stopping")
break
all_items.extend(items)
log.info(f"Got {len(items)} items (total: {len(all_items)})")
time.sleep(1)
return all_items
# --- Output ---
def save_results(items: list[dict], path: str):
with open(path, "w") as f:
json.dump(items, f, indent=2, ensure_ascii=False)
log.info(f"Saved {len(items)} items to {path}")
# --- Main ---
if __name__ == "__main__":
items = scrape_all(TARGET_URL)
save_results(items, OUTPUT_FILE)When building the scraper for the user, customize this template:
ITEM_SELECTOR, NAME_SELECTOR, PRICE_SELECTOR with real selectors from Phase 3| Scenario | API | Cost | Speed |
|---|---|---|---|
| Site has pre-built scraper (Amazon, LinkedIn, etc.) | Web Scraper API | Per record | Fast |
| Static HTML pages, no JS needed | Web Unlocker | Per request (success only) | Fast |
| Site exposes JSON API | Web Unlocker → API endpoint | Per request (success only) | Fastest |
| JS-rendered content (React, Vue, Angular) | Browser API | Per bandwidth | Medium |
| Infinite scroll | Browser API | Per bandwidth | Slow |
| Form submission / login required | Browser API | Per bandwidth | Medium |
| CAPTCHA-heavy sites | Browser API | Per bandwidth | Medium |
| Search engine results | SERP API | Per request | Fast |
Don't default to Browser API when Web Unlocker suffices. Browser API costs more (bandwidth-based) and is slower. Always try Web Unlocker first.
Don't use structural CSS selectors like div:nth-child(3) > span. They break when the site adds a banner or rearranges elements. Use data attributes or semantic selectors.
Don't hardcode pagination limits. Always check if the page returned actual items. An empty page means you've reached the end.
Don't skip the reconnaissance phase. Spending 2 minutes analyzing the HTML saves hours of debugging brittle selectors.
Don't forget error handling per item. One malformed product card shouldn't crash the entire scrape. Wrap individual item parsing in try/except.
Don't use networkidle with Browser API. SPAs never truly reach network idle. Use domcontentloaded + wait_for_selector instead.
Don't create a new browser session per page when scraping a list. If you're on a list page and clicking "next", you can stay in the same session. Only create new sessions for different base URLs.
Don't scrape URLs sequentially when you have many of them. Fetching 1,000+ URLs one-by-one with time.sleep(1) between each is unacceptably slow. Use concurrent requests with a semaphore. See references/concurrency-guide.md.
User says: "Build a scraper for Amazon product pages, I have a list of 200 ASINs"
Actions:
client.scrape.amazon.products_trigger() with batch of URLsResult: Complete Python script with async batch scraping, progress logging, JSON output.
User says: "I need to scrape all job listings from jobs.customsite.com including pagination"
Actions:
.job-card, .job-title, .company-name, .salary?page=N URL parameterResult: Complete Python script using Web Unlocker + BeautifulSoup with URL-based pagination.
User says: "Scrape product prices from a React SPA that loads data on scroll"
Actions:
div#root → client-renderedpage.evaluate() after content loadsResult: Async Playwright script with Browser API, infinite scroll handling, bandwidth optimization.
Cause: Site requires JavaScript rendering or has aggressive bot detection.
Solution: Escalate to Browser API. Also try adding data_format: "markdown" to see if the content is there but in a different format.
Cause: Site serves different HTML to different regions or user agents.
Solution: Add country parameter to Web Unlocker request to target the same region. Verify selectors on the actual HTML returned by the API, not browser DevTools.
Cause: Pagination logic is wrapping around or site uses inconsistent pagination. Solution: Track seen item IDs in a set. Break when duplicates appear. Verify the pagination URL pattern is correct.
Cause: Navigation timeout too short or site is slow to unblock.
Solution: Always set set_default_navigation_timeout(120_000). Use wait_until="domcontentloaded" not networkidle. Check if the site requires premium domains enabled on your zone.
Cause: Missing or invalid BRIGHTDATA_API_KEY environment variable.
Solution: Verify the key is set: echo $BRIGHTDATA_API_KEY. Get a fresh key from https://brightdata.com/cp/setting/users.
© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/scraper-builder of brightdata/skills.
Open the folder on GitHubat commit 81f51af
Scraper Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scraper Builder this skillbrightdata/skills | 264 | — | ~7.2k | Automated safety check: Pass | MIT | |
| Scrapedavila7/claude-code-templates | 32k | — | ~392 | Automated safety check: Pass | MIT | |
| Brightdatasundial-org/awesome-openclaw-skills | 663 | — | ~334 | Automated safety check: Pass | None | |
| BrightdataMicrock/ordinary-claude-skills | 401 | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| Tmuxtrpc-group/trpc-agent-go | 1.8k | 23 repos | ~868 | Automated safety check: Pass | Apache-2.0 | |
| Ketch1broseidon/ketch | 696 | 1 repos | ~3.9k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
Scrape any webpage as clean markdown via Bright Data Web Unlocker API.
sundial-org/awesome-openclaw-skills
Web scraping and search via Bright Data API. An agent skill from sundial-org/awesome-openclaw-skills.
Microck/ordinary-claude-skills
Progressive four-tier URL content scraping with automatic fallback strategy.
trpc-group/trpc-agent-go
Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.
1broseidon/ketch
Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
brightdata/skills
Replicate the visual style of any website and apply it to your existing codebase.
brightdata/skills
Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…
brightdata/skills
Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.
brightdata/skills
Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).
brightdata/skills
Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).
brightdata/skills
Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).
Works with
Categories
Build production-ready web scrapers for any website using Bright Data infrastructure. Scraper Builder is an agent skill from brightdata/skills. Build production-ready web scrapers for any website using Bright Data infrastructure.
Scraper Builder fits situations like: the user wants to build a scraper; create a crawler; extract data from a website; scrape product pages.
Run `npx skills add brightdata/skills --skill scraper-builder -a claude-code`. Or copy the skill folder (skills/scraper-builder in brightdata/skills) into .claude/skills/scraper-builder in your project. Claude Code loads it when a task matches its description.
Run `npx skills add brightdata/skills --skill scraper-builder -a codex`. Or copy the skill folder (skills/scraper-builder in brightdata/skills) into .agents/skills/scraper-builder in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill scraper-builder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scraper-builder, .gemini/skills/scraper-builder, .github/skills/scraper-builder and .opencode/skills/scraper-builder in your project.
Going by SKILL.md and its folder, Scraper Builder needs the command-line tools its instructions call (bash and curl) and credentials named BRIGHTDATA_API_KEY and API_KEY. Our summary lists: Python 3; Node.js; A credential in BRIGHTDATA_API_KEY; A credential in API_KEY.
SKILL.md names 5 domains. In commands or code: api.brightdata.com, amazon.com, target-site.com, docs.brightdata.com and brightdata.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scraper Builder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scraper Builder: Scrape (davila7/claude-code-templates, 32k stars), Brightdata (sundial-org/awesome-openclaw-skills, 663 stars), Brightdata (Microck/ordinary-claude-skills, 401 stars) and Tmux (trpc-group/trpc-agent-go, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.
Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.