Agent skill

Data Feeds

by brightdata in brightdata/skills

Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

MITAuto-check passedData & Analytics

Install Data Feeds

skills CLI
$ npx skills add brightdata/skills --skill data-feeds -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brightdata/skills data-feeds --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-feeds .claude/skills/data-feeds && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-feeds
GitHub stars
264
Token cost
~2.2k tokens
SKILL.md length
652 words
Files
4 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

  • Works in 6 steps: JSON parses cleanly: jq . returns 0 (or… → Record count matches expected. One URL… → No top-level error → …
  • The user wants clean JSON from a known platform URL rather than raw HTML
  • SKILL.md covers Setup gate (run first), Supported pipeline types…, Pick your path and Keyword- and multi-arg…, plus 4 more sections
  • Calls jq; reaches amazon.com and linkedin.com

What it does

Data Feeds is an agent skill from brightdata/skills. Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines). Use when the user wants clean JSON from a known platform URL rather than raw HTML. Hands off to scrape for unsupported URLs and to search when target URLs must be discovered first. Requires the Bright Data CLI; proactively guides install + login if missing.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/examples.md`, `references/flags.md` and `references/patterns.md`).

It sits in Data & Analytics, covering Web scraping and Schema markup. It works with Bright Data, Instagram, LinkedIn and TikTok. The licence is MIT.

When your agent uses it

  • The user wants clean JSON from a known platform URL rather than raw HTML
  • Tasks that involve Web scraping
  • Tasks that involve Schema markup

Example prompts

  • “/data-feeds”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. JSON parses cleanly: jq . returns 0 (or for --format ndjson, each line parses).
  2. Record count matches expected. One URL usually = one record, but reviews/posts/comments pipelines return arrays sized by what the platform…
  3. No top-level error
  4. No per-record error: for array results, ensure no record has an error field
  5. Core fields present for the pipeline type (examples)
  6. On failure: double --timeout and retry once. If still failing, bdata pipelines list to confirm the type name hasn't changed.

What it can do on your machine

Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • amazon.com
    • linkedin.com
    • instagram.com
    • maps.google.com
    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Feeds loads about 2.2k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 652 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 652 words, ~2,231 tokens.

Download SKILL.mdSave it as .claude/skills/data-feeds/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
data-feeds
description
Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (`bdata pipelines`). Use when the user wants clean JSON from a known platform URL rather than raw HTML. Hands off to `scrape` for unsupported URLs and to `search` when target URLs must be discovered first. Requires the Bright Data CLI; proactively guides install + login if missing.

Bright Data — Data Feeds (Pipelines)

Extract structured data from supported platforms via bdata pipelines. One call, clean JSON, no scraping logic. For unsupported URLs, hand off to scrape. To find target URLs first, hand off to search.

Setup gate (run first)

bash
if ! command -v bdata >/dev/null 2>&1; then
    echo "bdata CLI not installed — see bright-data-best-practices/references/cli-setup.md"
elif ! bdata zones >/dev/null 2>&1; then
    echo "bdata not authenticated — run: bdata login  (or: bdata login --device for SSH)"
fi

Halt and route to skills/bright-data-best-practices/references/cli-setup.md if either check fails.

Supported pipeline types (verified 2026-04-19)

Always verify with bdata pipelines list before hardcoding names — they change. Current 43 types:

amazon_product, amazon_product_reviews, amazon_product_search, apple_app_store, bestbuy_products, booking_hotel_listings, crunchbase_company, ebay_product, etsy_products, facebook_company_reviews, facebook_events, facebook_marketplace_listings, facebook_posts, github_repository_file, google_maps_reviews, google_play_store, google_shopping, homedepot_products, instagram_comments, instagram_posts, instagram_profiles, instagram_reels, linkedin_company_profile, linkedin_job_listings, linkedin_people_search, linkedin_person_profile, linkedin_posts, reddit_posts, reuter_news, tiktok_comments, tiktok_posts, tiktok_profiles, tiktok_shop, walmart_product, walmart_seller, x_posts, yahoo_finance_business, youtube_comments, youtube_profiles, youtube_videos, zara_products, zillow_properties_listing, zoominfo_company_profile

Naming note: inconsistent across platforms. amazon_product (singular), tiktok_profiles (plural), linkedin_person_profile (not linkedin_profile). Always copy from bdata pipelines list.

Pick your path

SituationAction
Know the platform + have URL(s)bdata pipelines <type> <url>
Don't know which pipeline fitsbdata pipelines list first
Pipeline takes keyword or multi-arg inputSee "Keyword- and multi-arg pipelines" below
Multiple URLs on the same pipeline typeshell loop with parallelism cap (see references/patterns.md)
Long job (reviews, company employees, big post feeds)raise --timeout 1800
URL is on an unsupported platformstop — hand off to scrape
Need to find URLs firsthand off to search

Keyword- and multi-arg pipelines (do NOT take a single URL)

A few pipelines take non-URL or multi-positional inputs. Invoke with no args to see the exact usage line from the CLI:

PipelineArgs
amazon_product_search<keyword> <domain_url> — e.g., "running shoes" https://www.amazon.com
linkedin_people_search<url> <first_name> <last_name> — search a company/school/URL for a named person
facebook_company_reviews<url> [num_reviews] — optional num_reviews defaults to 10
google_maps_reviews<url> [days_limit] — optional days_limit defaults to 3
youtube_comments<url> [num_comments] — optional num_comments defaults to 10

All other 37 pipelines take a single URL.

Action

Core commands:

bash
# List available pipeline types (source of truth)
bdata pipelines list

# Amazon product
bdata pipelines amazon_product \
    "https://www.amazon.com/dp/B08N5WRWNW" \
    --format json --pretty -o product.json

# Amazon product reviews (slower — reviews can be hundreds)
bdata pipelines amazon_product_reviews \
    "https://www.amazon.com/dp/B08N5WRWNW" \
    --timeout 1200 -o reviews.json

# Amazon product search (keyword + domain URL)
bdata pipelines amazon_product_search \
    "noise cancelling headphones" "https://www.amazon.com" \
    --format json --pretty -o search.json

# LinkedIn person profile
bdata pipelines linkedin_person_profile \
    "https://www.linkedin.com/in/example" -o person.json

# LinkedIn company
bdata pipelines linkedin_company_profile \
    "https://www.linkedin.com/company/example" -o company.json

# LinkedIn people search (url + first + last name)
bdata pipelines linkedin_people_search \
    "https://www.linkedin.com/company/example" "Jane" "Doe" \
    -o people.json

# Instagram posts
bdata pipelines instagram_posts \
    "https://www.instagram.com/example/" -o posts.json

# Google Maps reviews (url + days_limit, default 3)
bdata pipelines google_maps_reviews \
    "https://maps.google.com/?cid=1234567890" 90 -o reviews.json

# YouTube comments (url + num_comments, default 10)
bdata pipelines youtube_comments \
    "https://www.youtube.com/watch?v=abc123" 100 -o yt-comments.json

# NDJSON for big feeds (one record per line)
bdata pipelines linkedin_posts "https://www.linkedin.com/in/example" \
    --format ndjson -o posts.ndjson

# Raise polling timeout for long jobs
bdata pipelines amazon_product_reviews "<url>" --timeout 1800 -o out.json

Full flag reference + full type table: references/flags.md.

Verification gate

  1. JSON parses cleanly: jq . <output> returns 0 (or for --format ndjson, each line parses).

  2. Record count matches expected. One URL usually = one record, but reviews/posts/comments pipelines return arrays sized by what the platform shows. Always check:

    bash
    jq 'length' out.json                       # top-level array count
    # OR
    jq 'if type == "array" then length else 1 end' out.json
  3. No top-level error:

    bash
    jq -e 'if type == "object" then has("error") | not else true end' out.json \
        || { echo "pipeline reported error"; exit 1; }
  4. No per-record error: for array results, ensure no record has an error field:

    bash
    jq -e 'if type == "array" then map(has("error")) | any | not else true end' out.json \
        || echo "WARN: one or more records have error fields"

    Partial failures are silent — this check is non-optional.

  5. Core fields present for the pipeline type (examples):

    • amazon_product → .title + .price (or .final_price)
    • linkedin_person_profile → .name + .headline (or .position)
    • instagram_posts → .caption or .description + .url or .post_id
    • youtube_videos → .title + .video_id or .url

    Spot-check with jq keys on the first record to learn the exact schema.

  6. On failure: double --timeout and retry once. If still failing, bdata pipelines list to confirm the type name hasn't changed.

Show full SKILL.md (218 more words)Show less

Red flags

  • Using bdata scrape on Amazon/LinkedIn/TikTok/etc. when bdata pipelines <type> returns structured fields in one call. Loses structure and costs more time.
  • Looping bdata pipelines for large jobs without rate-limiting — each call can trigger a long-running pipeline on the server. Cap parallelism at 2–3.
  • Claiming success without the record-count + per-record error check. Partial failures are silent in pipeline output.
  • Hardcoding pipeline type names (amazon_products with an s, linkedin_profile without _person_, etc.) — they're inconsistent across platforms. Always copy from bdata pipelines list.
  • Using a tight --timeout on pipelines that legitimately take 5–15 minutes (reviews, company employees, big post feeds). Default 600s is a floor for small inputs; raise for long ones.
  • Calling a keyword- or multi-arg pipeline (amazon_product_search, linkedin_people_search, google_maps_reviews, facebook_company_reviews, youtube_comments) with URL-only args — will fail with "Usage: ...". Always check bdata pipelines <type> error output when in doubt.
  • Passing a pages_to_search third arg to amazon_product_search — it's hardcoded to 1 by the CLI and extra args are ignored.

References

  • references/flags.md — full pipelines flags + complete table of all 43 types with input shapes.
  • references/patterns.md — sync timeout tuning, shell-loop batching with parallelism cap, partial-failure detection, keyword-shaped pipeline cheatsheet, legacy curl fallback, shared verification checklist.
  • references/examples.md — (1) single Amazon product, (2) batch LinkedIn companies, (3) long reviews job with raised timeout, (4) mixed-platform workflow calling pipelines list first, (5) keyword-shaped amazon_product_search.

© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/data-feeds of brightdata/skills.

  • SKILL.md
  • references/examples.md
  • references/flags.md
  • references/patterns.md

Open the folder on GitHubat commit 81f51af

Compare with similar skills

Data Feeds next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Feeds compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Feeds this skillbrightdata/skills264—~2.2kAutomated safety check: PassMIT
Data Feedsdavila7/claude-code-templates32k—~1.7kAutomated safety check: PassMIT
Apify Multi-Platform Scraperapify/agent-skills2.4k2 repos~1.4kAutomated safety check: NotesNone
Scrapecreators APIScrapeCreators/social-media-research-skills3.3k1 repos~4kAutomated safety check: NotesMIT
Apify Content Analyticsmajiayu000/claude-skill-registry6663 repos~1.1kAutomated safety check: NotesMIT
Business Contact and Social Links Finderbrowser-act/skills6.1k1 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Data Feeds

    davila7/claude-code-templates

    Extract structured data from 40+ websites including Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, and more.

    32k GitHub stars~1.7k tokensUpdated today
    Marketing & SEOAuto-check passed
  • Official

    Scrapes public data from social, maps, search and review platforms by choosing from about a hundred Apify Actors and running them through the Apify CLI.

    2.4k GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check: notes
  • Scrapecreators API

    ScrapeCreators/social-media-research-skills

    Scrape and extract public data from 27+ social media platforms using the ScrapeCreators REST API.

    3.3k GitHub starsUsed in 1 repo~4k tokens
    Backend & APIsAuto-check: notes
  • Apify Content Analytics

    majiayu000/claude-skill-registry

    Track engagement metrics, measure campaign ROI, and analyze content performance across Instagram, Facebook, YouTube, and TikTok.

    666 GitHub starsUsed in 3 repos~1.1k tokens
    Data & AnalyticsAuto-check: notes
  • Finds a company's official website and social profiles from its name, or collects social links from a website URL, using BrowserAct templates run by a Python script.

    6.1k GitHub starsUsed in 1 repo~1.6k tokens
    Marketing & SEOAuto-check passed
  • Transcript Intelligence

    ScrapeCreators/social-media-research-skills

    A skill your agent uses when the user wants to summarize, analyze, or repurpose transcripts from TikTok, Instagram, YouTube, Facebook, X/Twitter, LinkedIn, Rumble, or Reddit video posts.

    3.3k GitHub stars~943 tokensUpdated 1 mo ago
    Marketing & SEOAuto-check: notes

More from brightdata/skills

All 14 skills in this repo
  • Design Mirror

    brightdata/skills

    Replicate the visual style of any website and apply it to your existing codebase.

    264 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Bright Data MCP

    brightdata/skills

    Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.

    264 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Live Research

    brightdata/skills

    Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).

    264 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Discover API

    brightdata/skills

    Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP).

    264 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Questions about Data Feeds

What does Data Feeds do?

Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines). Data Feeds is an agent skill from brightdata/skills. Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).

When should I use Data Feeds?

Data Feeds fits situations like: the user wants clean JSON from a known platform URL rather than raw HTML; tasks that involve Web scraping; tasks that involve Schema markup.

How do I install Data Feeds in Claude Code?

Run `npx skills add brightdata/skills --skill data-feeds -a claude-code`. Or copy the skill folder (skills/data-feeds in brightdata/skills) into .claude/skills/data-feeds in your project. Claude Code loads it when a task matches its description.

How do I install Data Feeds in Codex?

Run `npx skills add brightdata/skills --skill data-feeds -a codex`. Or copy the skill folder (skills/data-feeds in brightdata/skills) into .agents/skills/data-feeds in your project. Codex loads it when a task matches its description.

Can I use Data Feeds in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill data-feeds -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-feeds, .gemini/skills/data-feeds, .github/skills/data-feeds and .opencode/skills/data-feeds in your project.

What does Data Feeds need to run?

Going by SKILL.md and its folder, Data Feeds needs the command-line tools its instructions call (jq).

Does Data Feeds access the network?

SKILL.md names 5 domains. In commands or code: amazon.com, linkedin.com, instagram.com, maps.google.com and youtube.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Data Feeds safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Feeds use?

Data Feeds is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Feeds use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Data Feeds?

Skills that share tags, products or a category with Data Feeds: Data Feeds (davila7/claude-code-templates, 32k stars), Apify Multi-Platform Scraper (apify/agent-skills, 2.4k stars), Scrapecreators API (ScrapeCreators/social-media-research-skills, 3.3k stars) and Apify Content Analytics (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Feeds?

brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.