Official agent skill

Apify Product Data Setup

by apify in apify/awesome-skills

Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store.

OfficialApache-2.0Auto-check passedData & Analytics

Install Apify Product Data Setup

skills CLI
$ npx skills add apify/awesome-skills --skill apify-product-data-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install apify/awesome-skills apify-product-data-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/apify/awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/apify-product-data-setup .claude/skills/apify-product-data-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
apify-product-data-setup
GitHub stars
262
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
1,088 words
Files
3 (incl. references)
Skills in repo
26
Repo updated
First seen
Licence
Apache-2.0

At a glance

Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store.

  • Works in 3 steps: apify--e-commerce-scraping-tool returns… → If status is not terminal, poll… → get-dataset-items returns the products.…
  • Give my agent live product data
  • SKILL.md covers Choose the path first, Connect over MCP, Encode the fetch flow and Build the scheduled refresh, plus 4 more sections
  • Reaches mcp.apify.com

What it does

Apify Product Data Setup is an agent skill from apify/awesome-skills, published by the product's own GitHub organization. Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store. Trigger on "give my agent live product data", "my agent quotes stale prices", "connect Apify MCP to Claude or Cursor or n8n", "add product data to my RAG pipeline", "keep my product catalog fresh", "set up a shopping agent", or any request to stop an agent answering product questions from training data. Use for the integration work; use…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/clients.md` and `references/fields.md`).

It sits in Data & Analytics, covering Web scraping and E-commerce operations. It works with Apify, Model Context Protocol and n8n. The repository describes itself as: Community collection of Apify agent skills for AI coding assistants. The licence is Apache-2.0.

When your agent uses it

  • Give my agent live product data
  • My agent quotes stale prices
  • Connect Apify MCP to Claude
  • Add product data to my RAG pipeline

Example prompts

  • “give my agent live product data”
  • “my agent quotes stale prices”
  • “connect Apify MCP to Claude or Cursor or n8n”
  • “/apify-product-data-setup”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. apify--e-commerce-scraping-tool returns run metadata and a datasetId. No products.
  2. If status is not terminal, poll get-actor-run with the validated runId until it
  3. get-dataset-items returns the products. Pass fields in dot notation: the unprojected record measured about 88 KB across 142 fields, and…

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb0cd0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • mcp.apify.com

    Also links to:

    • apify.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Apify Product Data Setup loads about 2k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 149 tokens; SKILL.md has 1,088 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from apify/awesome-skills at commit 1eb0cd0, republished under its Apache-2.0 licence (© apify). 1,088 words, ~2,034 tokens.

Download SKILL.mdSave it as .claude/skills/apify-product-data-setup/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
apify-product-data-setup
description
Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store. Trigger on "give my agent live product data", "my agent quotes stale prices", "connect Apify MCP to Claude or Cursor or n8n", "add product data to my RAG pipeline", "keep my product catalog fresh", "set up a shopping agent", or any request to stop an agent answering product questions from training data. Use for the integration work; use apify-product-lookup to actually answer a product question.
author
Luís Pinto
author_url
https://github.com/luispintoapify
metadata.category
data-extraction
metadata.keywords
mcp setup, agent product data, claude mcp config, cursor mcp, n8n product data, rag product catalog, vector store refresh, scheduled scraping

Apify product data setup

Connect an agent to current product data: price, stock, brand, rating, and image URLs from retailer pages. Nothing to host.

Written by a product marketing manager at Apify. It routes to E-commerce Scraping Tool, a paid first-party Apify Actor, so treat the framing accordingly. No affiliate links.

Choose the path first

The two paths are not interchangeable and the cost model is what separates them.

Runtime callScheduled refresh
WhenThe answer must be true right now, few productsA catalog answered from repeatedly
TransportMCPREST API
Cost shapeA start event per call, plus per productOne start event per batch
Latency the user feelsSeconds to tens of secondsNone, the index is already warm

Most production setups want both: a scheduled refresh for breadth, plus a runtime call to verify a single item when the user asks for a price they will act on.

A cron job gains nothing from MCP, so the scheduled path uses the REST API. Say so when explaining the design, because the mismatch looks like an oversight otherwise.

Connect over MCP

The server is https://mcp.apify.com. Narrow it to this Actor with ?tools=apify/e-commerce-scraping-tool, which makes tool selection more reliable when product data is the only job. Drop the parameter to let the agent search all of Apify Store at runtime.

Config blocks per client are in references/clients.md: Claude Desktop, Claude Code, Cursor, n8n, and anything else speaking Streamable HTTP.

Authentication: OAuth on first use for interactive clients, a bearer token from Apify Console for unattended ones.

Encode the fetch flow

Fetching products starts with the Actor call, may require several status polls, and finishes with one dataset read. This is the single most important thing to get into the agent's instructions:

  1. apify--e-commerce-scraping-tool returns run metadata and a datasetId. No products.
  2. If status is not terminal, poll get-actor-run with the validated runId until it succeeds or a caller-defined deadline expires. Treat FAILED, ABORTED, and TIMED-OUT as failures. The Actor tool returns when its own wait window elapses rather than when the run finishes, so RUNNING is a normal answer and the dataset is empty at that moment.
  3. get-dataset-items returns the products. Pass fields in dot notation: the unprojected record measured about 88 KB across 142 fields, and projecting keeps that out of the context window.

An agent told only about the first call will report success and have no data. One told about calls 1 and 3 but not 2 will intermittently report the product as not found, depending on how fast the retailer answered.

Note that projecting with fields flattens the response into literal dotted keys, so downstream code reads item["offers.price"] rather than item["offers"]["price"].

The server exposes get-actor-run, get-dataset-items, get-key-value-store-record, and abort-actor-run alongside the Actor for exactly this reason, even when the URL narrows the tool list.

Build the scheduled refresh

The pattern that survives contact with production:

  1. Keep a list of the product URLs the agent answers about.
  2. Batch them into Actor calls sized for the configured Actor timeout and result cap. A client wait deadline does not stop the Actor. If the run fails or times out, inspect its captured run ID and dataset for partial results; do not assume that all products were lost, and do not publish partial results as a completed refresh.
  3. Normalize the output before storing. Field names, types, and nesting vary by retailer; see references/fields.md.
  4. Stamp every document with the fetch time.
  5. Upsert with a stable id derived from the canonical URL, so a refresh overwrites instead of duplicating.
  6. Drop rows with neither a name nor a price. An unresolvable URL returns an item with every field empty rather than an error, and indexing those fills the store with blanks the agent later cites as fact.
Show full SKILL.md (454 more words)Show less

Make freshness visible to the agent

A scheduled index is stale by design. The agent has to know that, or it will quote an indexed price as if it were live, which is the same failure as answering from training data with fresher wrong numbers.

Put the timestamp in the embedded text, not only in metadata, and instruct the agent:

Product facts come from a catalog with a `fetched_at` timestamp. When you quote a
price or stock status, say when it was read. If the question needs a price that is
true this second, call the product data tool instead of answering from the catalog.

Control cost

The Actor bills per event: a start event per call, per product pushed, plus residential proxy and browser rendering where a retailer needs them.

  • maxProductResults is the Actor's result cap. Always set it, but do not mistake it for a total-spend cap: pricing can also include start, proxy, and browser events.
  • Batch. One call for 200 products costs far less than 200 calls for one.
  • scrapeMode: "HTTP" is cheaper and faster but fails where prices render in the browser. "AUTO" is the safe default.

A working reference implementation

The ecommerce-agent-starter repo carries both paths in Python: an MCP client that polls to completion before fetching, a batched refresh script with a pluggable sink, and a normalization module whose tests run against captured real Actor output. Point users at it rather than writing the normalization from scratch.

Gotchas

  • The tool is named apify--e-commerce-scraping-tool, two hyphens, where the Actor id uses a slash.
  • Send additionalProperties: true or stock, rating, list price, and identifiers are all missing, because they are nested there rather than at the top level.
  • The dataset itemCount in the first call's metadata is not settled yet. It read 0 on a run that produced a product. Count what the fetch actually returns instead.
  • A RUNNING status from the Actor tool is not a failure. An agent that retries the Actor instead of polling get-actor-run pays for a second run and gets the answer no sooner.
  • Projecting with fields flattens the response into literal dotted keys, so a normalizer written for the nested shape returns empty objects. Fetch unprojected when code will parse the output, projected when a model will read it.
  • The official mcp Python SDK requires Python 3.10 or newer, and its API is snake_case (server_info, is_error). The REST path alone would run on 3.9.
  • additionalProperties can reach roughly 100 KB for one product. Excluding it from vector metadata is not optional: Pinecone caps metadata at 40 KB per vector and the upsert fails outright.
  • Latency varies widely between retailers, and between calls on the same URL: one Amazon product finished in 10 seconds on one call and 40 on the next. Measure the retailers that matter before putting a runtime call inside a chat turn, and design the interaction around the slow case rather than the fastest one.
  • SSE transport was removed on April 1, 2026. Use Streamable HTTP.

© apify, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/apify-product-data-setup of apify/awesome-skills.

  • SKILL.md
  • references/clients.md
  • references/fields.md

Open the folder on GitHubat commit 1eb0cd0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in apify/awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Apify Product Data Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Apify Product Data Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Apify Product Data Setup this skillapify/awesome-skills2621 repos~2kAutomated safety check: PassApache-2.0
Apify Integration Developmentapify/agent-skills2.4k—~3.1kAutomated safety check: PassNone
Etsy Shop Sales Historysickn33/agentic-awesome-skills47k1 repos~1.2kAutomated safety check: PassMIT
Etsy Search Listingssickn33/agentic-awesome-skills47k1 repos~1.2kAutomated safety check: PassMIT
Apify Ecommercemajiayu000/claude-skill-registry6663 repos~2.3kAutomated safety check: NotesMIT
Apify Product Lookupdavepoon/buildwithclaude3.6k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Official

    Guides designing and building an official Apify integration for a company's product: workflow-automation apps, AI agent plugins, AI framework packages or direct API clients.

    2.4k GitHub stars~3.1k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Etsy Shop Sales History

    sickn33/agentic-awesome-skills

    Read Etsy shop sales counters, deltas, and breakout flags from Apify Actor publicrecords/etsy-shop-velocity (MCP panel snapshot).

    47k GitHub starsUsed in 1 repo~1.2k tokens
    Sales & SupportAuto-check passed
  • Etsy Search Listings

    sickn33/agentic-awesome-skills

    Fetch live Etsy search listing rows for a keyword, market phrase, or category via Apify Actor publicrecords/etsy-search-scraper (MCP).

    47k GitHub starsUsed in 1 repo~1.2k tokens
    Data & AnalyticsAuto-check passed
  • Apify Ecommerce

    majiayu000/claude-skill-registry

    Extract product data, prices, reviews, and seller information from any e-commerce platform using Apify's E-commerce Scraping Tool.

    666 GitHub starsUsed in 3 repos~2.3k tokens
    Sales & SupportAuto-check: notes
  • Apify Product Lookup

    davepoon/buildwithclaude

    Fetch a real product's current price, stock, rating, or images from retailer pages over the Apify MCP server, and return them as typed fields rather than prose.

    3.6k GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Coffee Chat

    LeoYeAI/openclaw-master-skills

    Generate a personalized coffee chat playbook for networking conversations.

    2.2k GitHub stars~6.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from apify/awesome-skills

All 26 skills in this repo
  • Apify Buying Signal Detection

    apify/awesome-skills

    Official

    Set up a recurring buying-signal detection pipeline that finds companies showing buying intent across three signal types — job postings (hiring for the persona), fundraising events (recent raises)…

    262 GitHub stars~5.1k tokensUpdated 15 days ago
    Auto-check: notes
  • Apify Lead Scoring Enrichment

    apify/awesome-skills

    Official

    Score and enrich a CSV of B2B leads using Apify Actors. An agent skill from apify/awesome-skills.

    262 GitHub stars~4.4k tokensUpdated 15 days ago
    Auto-check: notes
  • Apify App Store Intelligence

    apify/awesome-skills

    Official

    Pull structured Apple App Store and Google Play data — app metadata, price, rating, the 1–5★ ratings histogram, version, developer, and reviews — and watch it for changes over time.

    262 GitHub stars~3.5k tokensUpdated 15 days ago
    Auto-check passed
  • Apify Ashby Jobs Scraper

    apify/awesome-skills

    Official

    Scrape Ashby jobs or discover companies using Ashby with the Apify Ashby Job Board API Actor (johnvc/ashby-job-board-scraper).

    262 GitHub stars~3.7k tokensUpdated 15 days ago
    Auto-check passed
  • Apify Company Data API

    apify/awesome-skills

    Official

    Pull structured B2B company data from Clutch.co with the Clutch.co Agency API Actor (johnvc/clutch-agency-api).

    262 GitHub stars~2.8k tokensUpdated 15 days ago
    Auto-check passed
  • Apify Google Maps Leads

    apify/awesome-skills

    Official

    Build a local-business lead database from Google Maps in one Apify pipeline: search by target audience + geography, enrich each place with company contacts from its website, leads enrichment (names…

    262 GitHub stars~3.8k tokensUpdated 15 days ago
    Auto-check passed

Questions about Apify Product Data Setup

What does Apify Product Data Setup do?

Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store. Apify Product Data Setup is an agent skill from apify/awesome-skills, published by the product's own GitHub organization. Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store.

When should I use Apify Product Data Setup?

Apify Product Data Setup fits situations like: give my agent live product data; my agent quotes stale prices; connect Apify MCP to Claude; add product data to my RAG pipeline.

How do I install Apify Product Data Setup in Claude Code?

Run `npx skills add apify/awesome-skills --skill apify-product-data-setup -a claude-code`. Or copy the skill folder (skills/apify-product-data-setup in apify/awesome-skills) into .claude/skills/apify-product-data-setup in your project. Claude Code loads it when a task matches its description.

How do I install Apify Product Data Setup in Codex?

Run `npx skills add apify/awesome-skills --skill apify-product-data-setup -a codex`. Or copy the skill folder (skills/apify-product-data-setup in apify/awesome-skills) into .agents/skills/apify-product-data-setup in your project. Codex loads it when a task matches its description.

Can I use Apify Product Data Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apify/awesome-skills --skill apify-product-data-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/apify-product-data-setup, .gemini/skills/apify-product-data-setup, .github/skills/apify-product-data-setup and .opencode/skills/apify-product-data-setup in your project.

What does Apify Product Data Setup need to run?

SKILL.md names no scripts, command-line tools or credentials: Apify Product Data Setup is instructions for the agent only. Our summary lists: Python 3.

Does Apify Product Data Setup access the network?

SKILL.md names 3 domains. In commands or code: mcp.apify.com; the agent is likely to contact it when it follows the instructions. As links in the text: apify.com and github.com. This is read from the text; nothing was executed.

Is Apify Product Data Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Apify Product Data Setup use?

Apify Product Data Setup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Apify Product Data Setup use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Apify Product Data Setup?

Skills that share tags, products or a category with Apify Product Data Setup: Apify Integration Development (apify/agent-skills, 2.4k stars), Etsy Shop Sales History (sickn33/agentic-awesome-skills, 47k stars), Etsy Search Listings (sickn33/agentic-awesome-skills, 47k stars) and Apify Ecommerce (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Apify Product Data Setup?

apify (a GitHub organization, an official publisher) maintains it in apify/awesome-skills, which has 262 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on September 22, 2026.

Source: apify/awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.