Agent skill

Scrape

by brightdata in brightdata/skills

Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape).

MITAuto-check passedData & Analytics

Install Scrape

skills CLI
$ npx skills add brightdata/skills --skill scrape -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brightdata/skills scrape --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scrape .claude/skills/scrape && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scrape
GitHub stars
264
Token cost
~1.2k tokens
SKILL.md length
417 words
Files
4 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape).

  • Works in 4 steps: Non-empty output: test -s "$out_path" —… → Not a block page — grep the output for… → Expected markers present for the task:… → …
  • The user wants to fetch a page
  • SKILL.md covers Setup gate (run first), Pick your path, Action and Verification gate (run before…, plus 2 more sections
  • Calls just

What it does

Scrape is an agent skill from brightdata/skills. Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape). Use when the user wants to fetch a page, extract content from a list of URLs, or crawl paginated listings. Hands off to data-feeds for supported platforms (Amazon, LinkedIn, TikTok, Instagram, YouTube, Reddit, etc.) and to search when URLs must be discovered first. Requires the Bright Data CLI; proactively guides install + login if missing.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/examples.md`, `references/flags.md` and `references/patterns.md`).

It sits in Data & Analytics, covering Web scraping. It works with Bright Data, LinkedIn, TikTok and Instagram. The licence is MIT.

When your agent uses it

  • The user wants to fetch a page
  • Extract content from a list of URLs
  • Crawl paginated listings

Example prompts

  • “/scrape”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Non-empty output: test -s "$out_path" — or, for stdout, at least 200 bytes of content.
  2. Not a block page — grep the output for any of these signatures (case-insensitive)
  3. Expected markers present for the task: e.g., a product page should contain a price pattern (\$\d); an article should contain at least one…
  4. On failure, escalation ladder

What it can do on your machine

Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scrape loads about 1.2k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 417 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 417 words, ~1,157 tokens.

Download SKILL.mdSave it as .claude/skills/scrape/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
scrape
description
Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (`bdata scrape`). Use when the user wants to fetch a page, extract content from a list of URLs, or crawl paginated listings. Hands off to `data-feeds` for supported platforms (Amazon, LinkedIn, TikTok, Instagram, YouTube, Reddit, etc.) and to `search` when URLs must be discovered first. Requires the Bright Data CLI; proactively guides install + login if missing.

Bright Data — Scrape

Get clean content (markdown, HTML, JSON, screenshot) from one or more URLs via the Bright Data CLI. This skill owns the "fetch raw or lightly-structured content" job. For platform-specific structured data (Amazon, LinkedIn, TikTok, etc.), stop and use data-feeds instead — you'll get clean JSON without selector logic.

Setup gate (run first)

Before any scrape, verify the CLI is installed and authenticated:

bash
if ! command -v bdata >/dev/null 2>&1; then
    echo "bdata CLI not installed — see bright-data-best-practices/references/cli-setup.md"
elif ! bdata zones >/dev/null 2>&1; then
    echo "bdata not authenticated — run: bdata login  (or: bdata login --device for SSH)"
fi

If either check fails, halt and route the user to skills/bright-data-best-practices/references/cli-setup.md. Do not attempt the legacy curl fallback silently — ask the user first.

Pick your path

SituationAction
Single URLbdata scrape <url> -f markdown
Small list (≤ ~20 URLs)shell loop, 1 at a time (see references/patterns.md)
Larger list (dozens+)xargs -P 4 with parallelism cap (see references/patterns.md)
Paginated listingscrape page 1 → extract next-page URL → append → repeat (see references/examples.md)
JS-heavy / login-gated / interaction-requiredescalate to bdata browser (see brightdata-cli skill)
Amazon, LinkedIn, TikTok, Instagram, YouTube, Reddit, …stop — hand off to data-feeds
No URL yet, just a topichand off to search

Action

Core commands:

bash
# Clean markdown (default)
bdata scrape "https://example.com/article" -f markdown -o article.md

# Raw HTML (when you need the DOM)
bdata scrape "https://example.com" -f html -o page.html

# Structured JSON (when the Unlocker returns parsed fields)
bdata scrape "https://example.com" -f json --pretty -o page.json

# Visual snapshot (saves PNG)
bdata scrape "https://example.com" -f screenshot -o page.png

# Geo-targeted (override the exit country)
bdata scrape "https://example.com" --country de -f markdown

Full flag reference: references/flags.md.

Verification gate (run before claiming success)

  1. Non-empty output: test -s "$out_path" — or, for stdout, at least 200 bytes of content.
  2. Not a block page — grep the output for any of these signatures (case-insensitive):
    • Access Denied
    • Just a moment
    • Attention Required
    • Checking your browser
    • captcha
    • cf-browser-verification
    • cloudflare (with < 2KB total body)
  3. Expected markers present for the task: e.g., a product page should contain a price pattern (\$\d); an article should contain at least one <h1> or # heading.
  4. On failure, escalation ladder:
    • Retry with a different --country (e.g., --country de if the origin site is US)
    • Escalate to bdata browser for full JS rendering (hand off to brightdata-cli skill)

Do not report success until all checks above pass.

Show full SKILL.md (125 more words)Show less

Red flags

  • Claiming success without inspecting the output.
  • Silencing errors with 2>/dev/null — you'll miss auth failures and rate-limit errors.
  • Running bdata scrape on Amazon/LinkedIn/TikTok/Instagram/YouTube/Reddit URLs — these are supported by data-feeds and return structured data directly. Scraping loses the structure.
  • Scraping the same URL repeatedly in the same task — cache the first result.
  • Looping bdata scrape sequentially for large lists instead of using xargs -P 4 (or similar) with a parallelism cap.
  • Using curl against api.brightdata.com directly — legacy path; only when the CLI isn't available.

References

  • references/flags.md — every flag with when-to-use notes.
  • references/patterns.md — shell-loop batching, xargs parallelism, pagination recipe, retry/backoff, block-page recovery chain, legacy curl fallback.
  • references/examples.md — (1) single page → markdown, (2) batch a list of URLs with parallelism cap, (3) paginated listing, (4) block-page recovery.

© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/scrape of brightdata/skills.

  • SKILL.md
  • references/examples.md
  • references/flags.md
  • references/patterns.md

Open the folder on GitHubat commit 81f51af

Compare with similar skills

Scrape next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scrape compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scrape this skillbrightdata/skills264—~1.2kAutomated safety check: PassMIT
Apify Multi-Platform Scraperapify/agent-skills2.4k2 repos~1.4kAutomated safety check: NotesNone
Scrapecreators APIScrapeCreators/social-media-research-skills3.3k1 repos~4kAutomated safety check: NotesMIT
Google Maps ScraperMahanaicoach/google-maps-scraper-kit1.3k—~2.8kAutomated safety check: PassMIT
Content Ideasbradautomates/content-ideas130—~5.4kAutomated safety check: NotesMIT
Business Contact and Social Links Finderbrowser-act/skills6.1k1 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Official

    Scrapes public data from social, maps, search and review platforms by choosing from about a hundred Apify Actors and running them through the Apify CLI.

    2.4k GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check: notes
  • Scrapecreators API

    ScrapeCreators/social-media-research-skills

    Scrape and extract public data from 27+ social media platforms using the ScrapeCreators REST API.

    3.3k GitHub starsUsed in 1 repo~4k tokens
    Backend & APIsAuto-check: notes
  • Google Maps Scraper

    Mahanaicoach/google-maps-scraper-kit

    Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.

    1.3k GitHub stars~2.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Content Ideas

    bradautomates/content-ideas

    Your For You page for content creators. An agent skill from bradautomates/content-ideas.

    130 GitHub stars~5.4k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check: notes
  • Finds a company's official website and social profiles from its name, or collects social links from a website URL, using BrowserAct templates run by a Python script.

    6.1k GitHub starsUsed in 1 repo~1.6k tokens
    Marketing & SEOAuto-check passed
  • Website Browsing Skills Index

    browsing-skills/browsing-skills

    Umbrella skill for a library of website-specific browsing skills. Use when the user's request targets one of these specific websites: <!-- DOMAINS:START…

    116 GitHub stars~1.6k tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed

More from brightdata/skills

All 14 skills in this repo
  • Design Mirror

    brightdata/skills

    Replicate the visual style of any website and apply it to your existing codebase.

    264 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Bright Data MCP

    brightdata/skills

    Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.

    264 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Live Research

    brightdata/skills

    Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).

    264 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Data Feeds

    brightdata/skills

    Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

    264 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Scrape

What does Scrape do?

Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape). Scrape is an agent skill from brightdata/skills. Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape).

When should I use Scrape?

Scrape fits situations like: the user wants to fetch a page; extract content from a list of URLs; crawl paginated listings.

How do I install Scrape in Claude Code?

Run `npx skills add brightdata/skills --skill scrape -a claude-code`. Or copy the skill folder (skills/scrape in brightdata/skills) into .claude/skills/scrape in your project. Claude Code loads it when a task matches its description.

How do I install Scrape in Codex?

Run `npx skills add brightdata/skills --skill scrape -a codex`. Or copy the skill folder (skills/scrape in brightdata/skills) into .agents/skills/scrape in your project. Codex loads it when a task matches its description.

Can I use Scrape in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill scrape -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scrape, .gemini/skills/scrape, .github/skills/scrape and .opencode/skills/scrape in your project.

What does Scrape need to run?

Going by SKILL.md and its folder, Scrape needs the command-line tools its instructions call (just).

Does Scrape access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scrape safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scrape use?

Scrape is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scrape use?

About 1.2k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Scrape?

Skills that share tags, products or a category with Scrape: Apify Multi-Platform Scraper (apify/agent-skills, 2.4k stars), Scrapecreators API (ScrapeCreators/social-media-research-skills, 3.3k stars), Google Maps Scraper (Mahanaicoach/google-maps-scraper-kit, 1.3k stars) and Content Ideas (bradautomates/content-ideas, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scrape?

brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.