Agent skill

Brightdata SDK JS

by brightdata in brightdata/skills

Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

MITAuto-check passedProductivity & Automation

Install Brightdata SDK JS

skills CLI
$ npx skills add brightdata/skills --skill brightdata-sdk-js -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brightdata/skills brightdata-sdk-js --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/js-sdk-best-practices .claude/skills/brightdata-sdk-js && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
brightdata-sdk-js
GitHub stars
264
Token cost
~3k tokens
SKILL.md length
851 words
Files
5 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

  • The user is working in Node.js/TypeScript and asks to scrape
  • SKILL.md covers Setup gate (do first), Service Selection (decide…, ⚠️ Key differences from the… and Method Names: Verify Before…, plus 6 more sections
  • Calls npm; reaches amazon.com and instagram.com; needs BRIGHTDATA_API_TOKEN
  • Find information from websites

What it does

Brightdata SDK JS is an agent skill from brightdata/skills. Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk). Use when the user is working in Node.js/TypeScript and asks to "scrape", "get data from", "extract", "search for", or "find" information from websites. Also use when the user mentions specific platforms like Amazon, LinkedIn, Instagram, Facebook, TikTok, YouTube, Reddit, Pinterest, ChatGPT, Perplexity, or DigiKey, or asks for "bulk data", "historical data", or "dataset" from JS. Covers scraping, SERP search, AI…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/advanced.md`, `references/datasets-overview.md` and `references/scrapers.md`).

It sits in Productivity & Automation, covering Web scraping and Web search. It works with Bright Data, Python, JavaScript and TypeScript. The licence is MIT.

When your agent uses it

  • The user is working in Node.js/TypeScript and asks to scrape
  • Find information from websites
  • The user mentions specific platforms like Amazon
  • Asks for bulk data

Example prompts

  • “scrape”
  • “get data from”
  • “extract”
  • “/brightdata-sdk-js”

Requirements

  • Python 3
  • Node.js
  • A credential in BRIGHTDATA_API_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • amazon.com
    • instagram.com

    Also links to:

    • brightdata.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BRIGHTDATA_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Brightdata SDK JS loads about 3k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 168 tokens; SKILL.md has 851 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 851 words, ~2,984 tokens.

Download SKILL.mdSave it as .claude/skills/brightdata-sdk-js/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
brightdata-sdk-js
description
Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (`@brightdata/sdk`). Use when the user is working in Node.js/TypeScript and asks to "scrape", "get data from", "extract", "search for", or "find" information from websites. Also use when the user mentions specific platforms like Amazon, LinkedIn, Instagram, Facebook, TikTok, YouTube, Reddit, Pinterest, ChatGPT, Perplexity, or DigiKey, or asks for "bulk data", "historical data", or "dataset" from JS. Covers scraping, SERP search, AI discovery, datasets, browser automation, and Scraper Studio. For Python, use brightdata-sdk; for the terminal CLI, use brightdata-cli.
metadata.author
brightdata
metadata.version
1.0
metadata.package
@brightdata/sdk
metadata.repository
https://github.com/brightdata/sdk-js

Bright Data JavaScript SDK

Access web data through a unified Node.js/TypeScript SDK (@brightdata/sdk). One client, several services: web unlocking (scrapeUrl), platform scraping (scrape.<platform>), SERP search (search.google/bing/yandex), AI discovery (discover), datasets, browser automation, and Scraper Studio.

Requires Node.js ≥ 20. Ships ESM + CommonJS with full TypeScript types. The client is exported as bdclient (lowercase — not a typo).

Setup gate (do first)

bash
npm install @brightdata/sdk     # or: pnpm add / yarn add

The client reads the BRIGHTDATA_API_TOKEN env var, or you pass { apiKey }. Get a token at https://brightdata.com/cp/setting/users.

javascript
// ESM (package.json "type": "module", or .mjs)
import { bdclient } from '@brightdata/sdk';

// CommonJS (.cjs / "type": "commonjs")
const { bdclient } = require('@brightdata/sdk');

const client = new bdclient();             // reads BRIGHTDATA_API_TOKEN
// const client = new bdclient({ apiKey: '...' });
try {
  const html = await client.scrapeUrl('https://example.com');
} finally {
  await client.close();                    // always close (or `await using`)
}

await using client = new bdclient() (TS 5.2+ / Node ≥20) auto-closes at scope end.

Service Selection (decide first, then look up the method)

Pick the service BEFORE reaching for a specific method. Most routing mistakes come from skipping this step and pattern-matching on keywords.

Have a URL?
  ├── On a supported platform (Amazon, LinkedIn, Facebook, Instagram, YouTube,
  │   TikTok, Reddit, Pinterest, ChatGPT, Perplexity, DigiKey)?
  │     → Platform scraping: client.scrape.<platform>.<method>(urls, opts?)
  │
  ├── Generic page (no dedicated platform scraper)?
  │     → Web unlocker: client.scrapeUrl(url, { dataFormat: 'markdown' })
  │
  └── Need login / JS / click-scroll-fill / CAPTCHA / multi-step nav?
        → Browser API: client.browser.getConnectUrl() (connect via Playwright)

No URL?
  ├── Want entities matching natural-language criteria
  │   ("find AI startups in Berlin", "competitors of Acme")?
  │     → Discover: client.discover(query, { intent })
  │
  ├── Want web pages / search-result links ("search Google for X")?
  │     → SERP: client.search.google(query) [or .bing / .yandex]
  │
  ├── Want to search WITHIN a platform ("Amazon products by keyword",
  │   "LinkedIn jobs", "Instagram reels by profile")?
  │     → Platform discovery: client.scrape.<platform>.discover*(filters)
  │       or amazon.productSearch(...)  (NOTE: client.search is SERP-only here)
  │
  └── Want bulk/historical data at scale?
        → Datasets: client.datasets.<name>.query(filter) → .download(snapshotId)

Edge cases:

  • Supported-platform URL BUT user mentions login/click/scroll/JS → Browser API (the interaction trumps the platform).
  • Supported-platform scrape returns 403/blocked → fall back to client.scrapeUrl() (web unlocker).
  • "Find/research who are X" with a URL alongside ("competitors of acme.com") → still Discover; the URL is context, not the scrape target.

⚠️ Key differences from the Python SDK

If you know the Python SDK (brightdata-sdk), the JS surface differs — do not port names blindly:

ConceptPythonJavaScript
ClientSyncBrightDataClient / BrightDataClientbdclient (single, all async)
Web unlockerclient.scrape_url(url=...)client.scrapeUrl(url, opts)
Platform searchclient.search.amazon.products(...)none — use scrape.amazon.productSearch / discover*
client.searchSERP and platform searchSERP only (google/bing/yandex)
Batch*_trigger methods + job.wait()pass a string[] to one call, or *Trigger + job.wait()
Namingsnake_casecamelCase
Datasetsclient.datasets.amazon_productsclient.datasets.amazonProducts

Method Names: Verify Before Asserting

Before claiming a platform method exists/doesn't exist, consult references/scrapers.md — it lists every platform's verified methods. The SDK ships TypeScript types, so in a typed project you can also let the compiler/editor confirm a method exists.

Each platform exposes up to three method styles (see references/scrapers.md):

  • collect<Thing>(input, opts?) — returns the rows directly (object[]).
  • <thing>(input, opts?) — orchestrated (trigger → poll → download); returns { data, status, rowCount }. Default to this.
  • discover<Thing>By<X>(filters, opts?) — find items by keyword/category/URL filter instead of by direct URL.

Likely hallucinations (do NOT write these — verify in references/scrapers.md):

WrongRight
client.search.amazon.products(...)client.scrape.amazon.productSearch(...) (no platform search router in JS)
client.scrape.linkedin.people(...)client.scrape.linkedin.profiles(urls)
client.scrape.chatgpt...client.scrape.chatGPT... (camelCase G+T)
client.datasets.amazon_productsclient.datasets.amazonProducts
BrightDataClient / new BdClient()bdclient (all lowercase)

Useful standalone methods

MethodWhat it does
client.scrapeUrl(url | url[], opts?)Web unlocker — any URL → html / markdown / json / screenshot. Pass an array for parallel batch.
client.search.google(q | q[], opts?)SERP results (also .bing, .yandex). Array = batch.
client.discover(query, { intent })AI-ranked entity/page discovery.
client.discoverTrigger(query, opts?)Non-blocking discover → Job with .wait() / .fetch().
client.datasets.list()List all available datasets at runtime.
client.scraperStudio.run(collectorId, { input })Run a custom Scraper Studio collector.
client.browser.getConnectUrl({ country })CDP WebSocket URL for Playwright/Puppeteer/Selenium.
client.listZones()List active Bright Data zones.
client.saveResults(data, { filename, format })Write results to a file.
client.close()Close HTTP connections. Always call when done.
Show full SKILL.md (410 more words)Show less

Gotchas

  • Always await client.close() (or await using). The client holds a keep-alive transport; leaking it hangs the process.
  • client.search is SERP-only (google/bing/yandex). To search within a platform, use that platform's discover* / productSearch scraper methods — there is no client.search.amazon.
  • chatGPT is camelCase on client.scrape — client.scrape.chatGPT.search(...), not .chatgpt.
  • Batch: pass an array, don't loop. scrapeUrl([...urls]) and search.google([...queries]) run in parallel internally. For platform scrapers with many inputs, prefer the orchestrated method with an array, or the *Trigger + job.wait() pattern (see references/advanced.md). Don't wrap blocking calls in Promise.all of single-item calls — you'll fight the rate limiter.
  • Orchestrated methods block for minutes. products/profiles/etc. trigger a job and poll until ready (often 2–10 min). Don't set tiny pollTimeouts; tune via { pollInterval, pollTimeout }. Defaults are calibrated.
  • Datasets are historical/bulk, not live. Need current data → use platform scrapers or scrapeUrl. query() returns a snapshotId; download() blocks until the snapshot is ready.
  • Web unlocker fallback on 403. If a platform scraper is blocked, retry the URL through client.scrapeUrl().
  • Don't add your own retry loop. The transport already retries network/timeout errors with backoff; double-retrying wastes credits. Tune timeout / rateLimit on the constructor instead.
  • Cost hierarchy (cheapest first): datasets → SERP → platform scrapers → web unlocker → discover → Scraper Studio → Browser API. Prefer the cheapest service that satisfies the request.

Error handling

All errors extend BRDError. Import the specific classes to branch:

javascript
import { bdclient, ValidationError, AuthenticationError, BRDError } from '@brightdata/sdk';

try {
  const data = await client.scrape.amazon.products(['https://www.amazon.com/dp/B0D77BX8Y4']);
} catch (err) {
  if (err instanceof AuthenticationError) { /* bad/expired token */ }
  else if (err instanceof ValidationError) { /* bad input */ }
  else if (err instanceof BRDError) { /* any SDK error */ }
  else throw err;
}

Classes: BRDError (base), ValidationError, AuthenticationError, ZoneError, NetworkError, NetworkTimeoutError, TimeoutError, APIError, DataNotReadyError, FSError.

Examples

Scrape a generic page as markdown
javascript
const md = await client.scrapeUrl('https://example.com/article', { dataFormat: 'markdown' });
Get an Amazon product (orchestrated — default)
javascript
const res = await client.scrape.amazon.products(['https://www.amazon.com/dp/B0D77BX8Y4']);
console.log(res.status, res.rowCount, res.data);
SERP, batched
javascript
const results = await client.search.google(['best running shoes', 'best trail shoes'], { country: 'us' });
Find entities (Discover)
javascript
const startups = await client.discover('AI startups in Berlin', {
  intent: 'early-stage machine-learning companies',
  numResults: 10,
});
Bulk/historical via datasets
javascript
const ds = client.datasets;
const snapshotId = await ds.instagramProfiles.query({ url: 'https://www.instagram.com/natgeo/' }, { records_limit: 50 });
// download() blocks until the snapshot is ready
const rows = await ds.instagramProfiles.download(snapshotId);

Troubleshooting

  • AuthenticationError / 401 — token missing or invalid. Check BRIGHTDATA_API_TOKEN or the apiKey option.
  • 403 / blocked — site blocked the scraper. Retry the URL via client.scrapeUrl() (web unlocker).
  • Timeout — increase timeout (constructor, 1000–300000 ms) or the orchestrated pollTimeout; do not lower it.
  • "Dataset not found" — call client.datasets.list(); dataset names are camelCase (amazonProducts, linkedinProfiles).
  • Process hangs after work finishes — you forgot await client.close().
  • SSL/proxy errors in sandboxes — see references/advanced.md for zone/SSL options.
  • Rate-limit errors — set { rateLimit, ratePeriod } on the constructor, or batch via arrays / the trigger pattern instead of Promise.all.

When to load references

  • references/scrapers.md — when the user names Amazon, LinkedIn, Facebook, Instagram, YouTube, TikTok, Reddit, Pinterest, ChatGPT, Perplexity, DigiKey — verified per-platform method tables (collect / orchestrated / discover) and the scrapeUrl web-unlocker options.
  • references/search.md — SERP engines (search.google/bing/yandex) and the Discover API (discover / discoverTrigger), with all options.
  • references/datasets-overview.md — dataset names, the query → getStatus → download lifecycle, and filters.
  • references/advanced.md — constructor options & env vars, batch/trigger orchestration, Browser API (Playwright/Puppeteer), Scraper Studio, error classes, and zones/SSL.

© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/js-sdk-best-practices of brightdata/skills.

  • SKILL.md
  • references/advanced.md
  • references/datasets-overview.md
  • references/scrapers.md
  • references/search.md

Open the folder on GitHubat commit 81f51af

Compare with similar skills

Brightdata SDK JS next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Brightdata SDK JS compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Brightdata SDK JS this skillbrightdata/skills264—~3kAutomated safety check: PassMIT
Scrapingbee CLIScrapingBee/scrapingbee-cli108—~3.2kAutomated safety check: NotesMIT
Skyvern Browser AutomationSkyvern-AI/skyvern23k—~1.9kAutomated safety check: PassAGPL-3.0
Tavily Search API Integrationandrewyng/context-hub14k—~1.1kAutomated safety check: PassMIT
Gemini Interactions APIAyuilos/Miffan192—~4.6kAutomated safety check: PassAGPL-3.0
Agent Readiness Auditindranilbanerjee/digital-marketing-pro8551 repos~3.9kAutomated safety check: PassMIT

Similar skills

  • Scrapingbee CLI

    ScrapingBee/scrapingbee-cli

    Fetch and read any web page, search the web, crawl a site, or pull structured data out of pages.

    108 GitHub stars~3.2k tokensUpdated yesterday
    Productivity & AutomationAuto-check: notes
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Tavily Search API Integration

    andrewyng/context-hub

    Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.

    14k GitHub stars~1.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    192 GitHub stars~4.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended)…

    855 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Openai Security Best Practices

    trailofbits/skills-curated

    Official

    Perform language and framework specific security best-practice reviews and suggest improvements.

    512 GitHub starsUsed in 9 repos~2.2k tokens
    SecurityAuto-check: notes

More from brightdata/skills

All 14 skills in this repo
  • Design Mirror

    brightdata/skills

    Replicate the visual style of any website and apply it to your existing codebase.

    264 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Bright Data MCP

    brightdata/skills

    Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.

    264 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Live Research

    brightdata/skills

    Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).

    264 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Data Feeds

    brightdata/skills

    Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

    264 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Discover API

    brightdata/skills

    Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP).

    264 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Questions about Brightdata SDK JS

What does Brightdata SDK JS do?

Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk). Brightdata SDK JS is an agent skill from brightdata/skills. Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

When should I use Brightdata SDK JS?

Brightdata SDK JS fits situations like: the user is working in Node.js/TypeScript and asks to scrape; find information from websites; the user mentions specific platforms like Amazon; asks for bulk data.

How do I install Brightdata SDK JS in Claude Code?

Run `npx skills add brightdata/skills --skill brightdata-sdk-js -a claude-code`. Or copy the skill folder (skills/js-sdk-best-practices in brightdata/skills) into .claude/skills/brightdata-sdk-js in your project. Claude Code loads it when a task matches its description.

How do I install Brightdata SDK JS in Codex?

Run `npx skills add brightdata/skills --skill brightdata-sdk-js -a codex`. Or copy the skill folder (skills/js-sdk-best-practices in brightdata/skills) into .agents/skills/brightdata-sdk-js in your project. Codex loads it when a task matches its description.

Can I use Brightdata SDK JS in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill brightdata-sdk-js -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/brightdata-sdk-js, .gemini/skills/brightdata-sdk-js, .github/skills/brightdata-sdk-js and .opencode/skills/brightdata-sdk-js in your project.

What does Brightdata SDK JS need to run?

Going by SKILL.md and its folder, Brightdata SDK JS needs the command-line tools its instructions call (npm) and credentials named BRIGHTDATA_API_TOKEN. Our summary lists: Python 3; Node.js; A credential in BRIGHTDATA_API_TOKEN.

Does Brightdata SDK JS access the network?

SKILL.md names 3 domains. In commands or code: amazon.com and instagram.com; the agent is likely to contact these when it follows the instructions. As links in the text: brightdata.com. This is read from the text; nothing was executed.

Is Brightdata SDK JS safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Brightdata SDK JS use?

Brightdata SDK JS is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Brightdata SDK JS use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.

What are the alternatives to Brightdata SDK JS?

Skills that share tags, products or a category with Brightdata SDK JS: Scrapingbee CLI (ScrapingBee/scrapingbee-cli, 108 stars), Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars), Tavily Search API Integration (andrewyng/context-hub, 14k stars) and Gemini Interactions API (Ayuilos/Miffan, 192 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Brightdata SDK JS?

brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.

Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.