Agent skill

Discover API

by brightdata in brightdata/skills

Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP).

MITAuto-check passedProductivity & Automation

Install Discover API

skills CLI
$ npx skills add brightdata/skills --skill discover-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brightdata/skills discover-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/discover-api .claude/skills/discover-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
discover-api
GitHub stars
264
Token cost
~2.4k tokens
SKILL.md length
779 words
Files
2 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP).

  • Works in 3 steps: Trigger a job → you get a task_id. → Poll with the task_id until status is… → Read results[].
  • A discovery job and retrieve ranked results (link
  • SKILL.md covers How it works (async: trigger →…, Pick your surface, Parameters (REST body —… and Modes (choose by goal), plus 5 more sections
  • Calls jq, curl and just; reaches api.brightdata.com; needs BRIGHTDATA_API_TOKEN

What it does

Discover API is an agent skill from brightdata/skills. Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Trigger a discovery job and retrieve ranked results (link, title, description, relevancescore) with optional parsed page content. Use when the user wants semantic/intent-based web search, "find pages about <topic that match <goal", web-grounded retrieval for an LLM, or results filtered by relevance rather than raw keyword rank. Covers the REST API (POST/GET /discover), the CLI (bdata discover), and the…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api-reference.md`).

It sits in Productivity & Automation, covering Web search, Retrieval-augmented generation and REST APIs. It works with Bright Data and Python. The licence is MIT.

When your agent uses it

  • A discovery job and retrieve ranked results (link
  • Relevancescore) with optional parsed page content
  • The user wants semantic/intent-based web search
  • Find pages about <topic that match <goal

Example prompts

  • “find pages about <topic that match <goal”
  • “/discover-api”

Requirements

  • Python 3
  • A credential in BRIGHTDATA_API_TOKEN

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Trigger a job → you get a task_id.
  2. Poll with the task_id until status is "done" (intermediate: "processing").
  3. Read results[].

What it can do on your machine

Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • curl
    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.brightdata.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BRIGHTDATA_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Discover API loads about 2.4k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 192 tokens; SKILL.md has 779 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~192
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 779 words, ~2,411 tokens.

Download SKILL.mdSave it as .claude/skills/discover-api/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
discover-api
description
Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Trigger a discovery job and retrieve ranked results (link, title, description, relevance_score) with optional parsed page content. Use when the user wants semantic/intent-based web search, "find pages about <topic> that match <goal>", web-grounded retrieval for an LLM, or results filtered by relevance rather than raw keyword rank. Covers the REST API (POST/GET /discover), the CLI (`bdata discover`), and the Python/JS SDKs (`client.discover`), including the standard/zeroRanking/deep/fast modes. This is the foundation skill for `live-research` and `rag-pipeline`. For keyword SERP use `search`; for structured platform data use `data-feeds`.
metadata.author
Bright Data
metadata.version
1.0
metadata.documentation
https://docs.brightdata.com/api-reference/discover/overview

Bright Data — Discover API

Discover is intent-ranked semantic web search. You give it a query plus an intent, and it returns results scored by AI relevance — optionally with the full parsed page content. It is the right primitive when result quality/relevance matters more than raw keyword rank, and the building block for retrieval (RAG), research, and knowledge-base pipelines.

Discover vs. the neighbors:

  • Keyword "what ranks for X" SERP → use the search skill (bdata search).
  • Structured data from a known platform (Amazon/LinkedIn/…) → use data-feeds.
  • A whole research brief or a RAG/search pipeline on top of Discover → use live-research or rag-pipeline (both call this API).

How it works (async: trigger → poll)

  1. Trigger a job → you get a task_id.
  2. Poll with the task_id until status is "done" (intermediate: "processing").
  3. Read results[].

The CLI and SDKs do the trigger+poll for you; the raw REST flow is shown below for when you need parameters the wrappers don't expose (notably mode).

Pick your surface

You are…Use
In a terminal, one-off or scriptedCLI: bdata discover
Writing Node/TS codeJS SDK: client.discover() — see js-sdk-best-practices
Writing Python codePython SDK: client.discover() — see python-sdk-best-practices
Need mode (deep/fast/zeroRanking) or include_imagesRaw REST (wrappers don't expose these yet)
CLI — bdata discover

Setup gate first:

bash
command -v bdata >/dev/null 2>&1 || echo "CLI missing — see bright-data-best-practices/references/cli-setup.md"
bdata zones >/dev/null 2>&1 || echo "not authenticated — run: bdata login"
bash
# Intent-ranked discovery, JSON
bdata discover "enterprise LLM platforms" \
  --intent "vendor pages with pricing" \
  --num-results 15 --json --pretty

# With parsed page content in one pass (for RAG / research)
bdata discover "webhook retry best practices" \
  --include-content --num-results 10 -o results.json

# Date-bounded
bdata discover "react server components" \
  --start-date 2025-01-01 --end-date 2025-12-31 --num-results 20 --json

Results live at .results[]; each has title, link, description, relevance_score, and content when --include-content. Full CLI flag list: search skill → references/flags.md.

SDK (one line each)
javascript
// JS — see js-sdk-best-practices for all options.
// VERIFIED v1.1.0: discover() returns a WRAPPER { success, data:[...], totalResults, cost, taskId, ... }
const res = await client.discover('Tesla battery tech', { intent: 'EV battery breakthroughs', numResults: 10, includeContent: true });
const rows = res.data;   // ← rows are in .data (NOT a bare array, NOT .results)
python
# Python — see python-sdk-best-practices (confirm whether rows come back directly, under .data, or .results)
out = client.discover(query="Tesla battery tech", intent="EV battery breakthroughs")
Raw REST (full control, incl. mode)
bash
# 1) Trigger
task_id=$(curl -s -X POST https://api.brightdata.com/discover \
  -H "Authorization: Bearer $BRIGHTDATA_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"query":"post-quantum cryptography adoption","intent":"enterprise migration guides","mode":"deep","num_results":20,"include_content":true}' \
  | jq -r '.task_id')

# 2) Poll until done
while :; do
  resp=$(curl -s "https://api.brightdata.com/discover?task_id=$task_id" -H "Authorization: Bearer $BRIGHTDATA_API_TOKEN")
  [ "$(echo "$resp" | jq -r '.status')" = "done" ] && break
  sleep 3
done
echo "$resp" | jq '.results'

Parameters (REST body — authoritative superset)

query is required; everything else is optional. The CLI/SDK expose a subset (see references/api-reference.md for the exact per-surface matrix).

ParamTypeDefaultNotes
querystring—required, ≤ 1500 chars
intentstring—goal descriptor, ≤ 3000 chars; strongly recommended — drives ranking
modeenumstandardstandard | zeroRanking | deep | fast (REST-only)
num_resultsint—1–20; ignored in zeroRanking
filter_keywordsstring[]—exact keywords that must appear
include_contentboolfalseparsed page/PDF content (PDF ≤ 50 MB, 30s); unsupported in zeroRanking
include_imagesboolfalseimage array (REST-only)
formatenumjsonjson | md (SDK accepts only json)
countrystringUS2-letter ISO
citystring—SERP city targeting
languagestringen31 languages
start_date / end_datestring—YYYY-MM-DD (REST-only)
remove_duplicatesbooltruededupe results (REST-only)

Modes (choose by goal)

ModeWhat it doesUse for
standard (default)balanced depth + AI rankinggeneral intent search
deepexhaustive, broader search; slowerlive-research, comprehensive topic coverage
fastoptimized for low latencytime-sensitive / interactive
zeroRankingno AI ranking, max raw volume; ignores num_results, no include_contentbulk corpus collection for rag-pipeline

mode is currently REST-only — the CLI and SDKs don't expose it. For deep coverage via the CLI/SDK, approximate with a high num_results + a sharp intent; for true deep/zeroRanking, use the raw REST flow above.

Show full SKILL.md (338 more words)Show less

Result shape

REST + CLI (verified) — rows live under results:

json
{
  "status": "done",
  "duration_seconds": 12.4,
  "timestamp": "2026-06-08T08:36:55.709Z",
  "results": [
    { "link": "https://…", "title": "…", "description": "…",
      "relevance_score": 0.87, "content": "…(when --include-content)…" }
  ]
}

JS SDK (verified v1.1.0) — rows live under data, inside a wrapper:

js
{ success: true, data: [ {link, title, description, relevance_score, content?} ],
  totalResults, cost, taskId, query, intent, durationSeconds, triggerSentAt, dataFetchedAt }

⚠️ Cross-surface gotcha: REST/CLI return rows under .results; the JS SDK returns them under .data (and wraps everything in {success, ...} — check success before reading data). Don't assume one shape across surfaces.

relevance_score is a float (snake_case). Higher = more relevant to intent. content is plain text by default (Markdown when REST format=md). A high relevance_score does not guarantee good content — pages can be 404 stubs or nav-only; gate on content length + "not found"/block-page signatures before use.

Verification gate

  1. Trigger returned a task_id (REST: status:"ok"). No id → check auth / that Discover is enabled on the account (403 if disabled).
  2. Polled to status:"done" before reading — never read results while processing.
  3. results[] non-empty — if empty, the query/intent is too narrow; loosen and retry. Don't claim success on empty.
  4. include_content bodies aren't block pages — grep content for captcha, Just a moment, Access Denied, cf-browser-verification (same list as scrape). Drop poisoned rows.
  5. Relevance sanity — if top relevance_scores are low or off-topic, sharpen intent (not just query).

Red flags

  • Passing only query with no intent — you lose the whole point (intent ranking). Always give an intent.
  • Treating Discover as keyword SERP — for "what ranks for X", use search.
  • Setting num_results > 20 — capped at 20; for more, run multiple targeted queries and dedup (see live-research).
  • Using zeroRanking then expecting include_content or num_results to apply — they don't.
  • Reading results before status:"done".
  • Treating an intermittent success:false (SDK) / empty results (REST/CLI) as a hard error — it's often transient; retry once with backoff first (see references/api-reference.md → Transient failures).
  • Fabricating relevance_score or content when a call fails — report the failure instead.

References

  • references/api-reference.md — full REST endpoint spec, the exact param matrix per surface (REST vs CLI vs JS SDK vs Python SDK), error codes, and limits.
  • live-research — multi-query Discover → dedup → synthesized cited brief.
  • rag-pipeline — Discover (include_content) → chunk → embed → retrieve.
  • search — keyword SERP (bdata search) when you don't need intent ranking.
  • js-sdk-best-practices / python-sdk-best-practices — client.discover() option details.

© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/discover-api of brightdata/skills.

  • SKILL.md
  • references/api-reference.md

Open the folder on GitHubat commit 81f51af

Compare with similar skills

Discover API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Discover API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Discover API this skillbrightdata/skills264—~2.4kAutomated safety check: PassMIT
Tavily Search API Integrationandrewyng/context-hub14k—~1.1kAutomated safety check: PassMIT
Brightdata Local Searchdavila7/claude-code-templates32k—~1.2kAutomated safety check: NotesMIT
DBoracle/skills872—~1.4kAutomated safety check: PassUPL-1.0
Proxy6VKirill/claude-lane-stack122—~3.6kAutomated safety check: PassMIT
Duckduckgo SearchAlexAI-MCP/hermes-CCC135—~1.6kAutomated safety check: PassMIT

Similar skills

  • Tavily Search API Integration

    andrewyng/context-hub

    Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.

    14k GitHub stars~1.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Brightdata Local Search

    davila7/claude-code-templates

    Set up and run local web searches using Bright Data SERP API with the unfancy-search pipeline (query expansion, SERP retrieval, RRF reranking).

    32k GitHub stars~1.2k tokensUpdated today
    Productivity & AutomationAuto-check: notes
  • DB

    oracle/skills

    Official

    Oracle Database guidance for SQL, PL/SQL, SQLcl, ORDS, Oracle Vector SDK, administration, app development, performance, security, migrations, and agent-safe database workflows.

    872 GitHub stars~1.4k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Proxy6

    VKirill/claude-lane-stack

    [RU: интеграция proxy6.net — покупка/продление прокси, пул, ipauth, scraping] proxy6.net REST API — RU proxy provider for IPv4/IPv4 Shared/IPv6/MTproto.

    122 GitHub stars~3.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Duckduckgo Search

    AlexAI-MCP/hermes-CCC

    Free web search via DuckDuckGo - no API key needed. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Productivity & AutomationAuto-check passed
  • Aivis

    onism1767-creator/potato

    Claude API web-search 可见度(代理测量):用冻结题库反复探测 Claude API(websearch), 以确定性规则统计品牌被提及/被引用的频率与结构,输出带不确定区间、可审计的报告。

    166 GitHub stars~757 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from brightdata/skills

All 14 skills in this repo
  • Design Mirror

    brightdata/skills

    Replicate the visual style of any website and apply it to your existing codebase.

    264 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Bright Data MCP

    brightdata/skills

    Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.

    264 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Live Research

    brightdata/skills

    Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).

    264 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Data Feeds

    brightdata/skills

    Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

    264 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Discover API

What does Discover API do?

Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Discover API is an agent skill from brightdata/skills. Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP).

When should I use Discover API?

Discover API fits situations like: A discovery job and retrieve ranked results (link; relevancescore) with optional parsed page content; the user wants semantic/intent-based web search; find pages about <topic that match <goal.

How do I install Discover API in Claude Code?

Run `npx skills add brightdata/skills --skill discover-api -a claude-code`. Or copy the skill folder (skills/discover-api in brightdata/skills) into .claude/skills/discover-api in your project. Claude Code loads it when a task matches its description.

How do I install Discover API in Codex?

Run `npx skills add brightdata/skills --skill discover-api -a codex`. Or copy the skill folder (skills/discover-api in brightdata/skills) into .agents/skills/discover-api in your project. Codex loads it when a task matches its description.

Can I use Discover API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill discover-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/discover-api, .gemini/skills/discover-api, .github/skills/discover-api and .opencode/skills/discover-api in your project.

What does Discover API need to run?

Going by SKILL.md and its folder, Discover API needs the command-line tools its instructions call (jq, curl and just) and credentials named BRIGHTDATA_API_TOKEN. Our summary lists: Python 3; A credential in BRIGHTDATA_API_TOKEN.

Does Discover API access the network?

SKILL.md names 1 domain. In commands or code: api.brightdata.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Discover API safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Discover API use?

Discover API is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Discover API use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Discover API?

Skills that share tags, products or a category with Discover API: Tavily Search API Integration (andrewyng/context-hub, 14k stars), Brightdata Local Search (davila7/claude-code-templates, 32k stars), DB (oracle/skills, 872 stars) and Proxy6 (VKirill/claude-lane-stack, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Discover API?

brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.