Agent skill

Data Scraper

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…

MITAuto-check passedData & Analytics

Install Data Scraper

skills CLI
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness data-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-scraper .claude/skills/data-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-scraper
GitHub stars
156
Token cost
~3.3k tokens
SKILL.md length
1,708 words
Files
7 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…

  • Data lives on a website with no usable API — listings
  • SKILL.md covers The legal gate — run this…, Extraction-path decision table, Tool picker and Robust selectors, plus 3 more sections
  • Runs Shell scripts from its folder
  • Public records — and the scrape must stay legal and not get blocked: legal gate

What it does

Data Scraper is an agent skill from ericrisco/rsc-harness. Use when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable selectors, pacing, resilience. NOT parsing bytes you already hold into fields (that is structured-extraction), NOT a documented API or key (that is api-connector-builder).

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/anti-bot.md`).

It sits in Data & Analytics, covering Web scraping and Document parsing. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Data lives on a website with no usable API — listings
  • Public records — and the scrape must stay legal and not get blocked: legal gate
  • Extraction path
  • Durable selectors

Example prompts

  • “/data-scraper”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Scraper loads about 3.3k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 1,708 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,708 words, ~3,310 tokens.

Download SKILL.mdSave it as .claude/skills/data-scraper/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
data-scraper
description
Use when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable selectors, pacing, resilience. NOT parsing bytes you already hold into fields (that is structured-extraction), NOT a documented API or key (that is api-connector-builder).
tags
web-scraping, crawler, anti-bot, robots-txt, gdpr-compliance, playwright, rate-limiting
recommends
api-connector-builder, structured-extraction, data-cleaning, gdpr-privacy, webhooks
origin
risco

Data scraper

You get bytes off websites you do not control — legally, and without getting blocked. You pick the cheapest extraction path that works, build selectors that survive a redesign, pace requests so the host neither bans nor sues you, and you write down the legal basis before the first request goes out.

One rule above all the others: scraping is the fallback, not the default. It is what you reach for only when no API serves the data. If the site has a documented API or you hold a key, stop — that is ../api-connector-builder/SKILL.md. And once you have the bytes, parsing them into fields is ../structured-extraction/SKILL.md, normalizing the rows is ../data-cleaning/SKILL.md. This skill ends the moment you hold the bytes.

In 2025-2026 scrapers do not fail on parsing. They fail on a terms-of-service breach, on GDPR exposure, or on being blocked after hammering a host. So the work runs in this order: legal gate → extraction path → tool → selectors → politeness → resilience. Skipping the gate is how you end up in Meta v. Bright Data.

Walk every item. Each ends in proceed, proceed-narrowed, or stop. One item at stop means the whole scrape stops until you resolve it. Depth and the case law are in references/legal-compliance.md.

The gates below are the ones with real legal exposure — contract, data-protection law, and anti-circumvention. robots.txt is not one of them: it is a voluntary convention with no statutory force, so it informs the decision and never blocks it on its own.

  • Is there an API? If yes, you are in the wrong skill — never scrape what an API serves. → otherwise proceed.
  • Did you read the ToS, and does login/auth apply? Scraping is most exposed as breach of contract when you accepted terms — typically by logging in (Meta v. Bright Data, 2024). Logged-out public data weakens that claim. Prefer logged-out public pages; never bypass auth. → narrow to public, or stop.
  • Is it personal data? Names, emails, photos, reviews, IP addresses all count. Scraping public personal data for a new purpose — aggregation, resale, AI training — is a severe GDPR breach with fines into the tens of millions of EUR. You need a lawful basis (usually legitimate interest) and data minimization. Filter out special categories at the source. → proceed-narrowed (basis + minimization, see ../gdpr-privacy/SKILL.md), or stop.
  • robots.txt and ai.txt? Read them — they are a convention, not law, so this item never stops a scrape on its own. A Disallow tells you the host would rather you did not, and Crawl-delay tells you its tolerance; both are useful intelligence about where you are likely to get blocked. Weigh it: honoring robots is the low-friction default and the cleanest evidence of good faith, but scraping your own site, one you have permission for, or public pages for research is legitimate whether or not robots allows it. → advisory: note the decision and move on.
  • Are you about to bypass a control you were shown? A CAPTCHA, a hard block, an enforced rate limit. Reddit v. Perplexity AI (2025) turns precisely on whether anti-bot measures were circumvented — a materially worse position than respectfully pacing public pages. Pacing public data is defensible; defeating a control is the frontier where you lose. → stop. Do not solve the CAPTCHA. Do not bypass the block.

hiQ v. LinkedIn established that scraping public data is not automatically a CFAA violation — but that is the floor, not a license. Public + logged-out + no-personal-data is the defensible quadrant, and honoring robots keeps it tidy without being what makes it lawful. Anything outside it, document why and get a human to sign off.

Extraction-path decision table

Walk down only when the rung above is unavailable. ~94% of modern sites are client-side rendered, yet most still ship machine-readable structured data — parse that before you launch a browser.

PathUse whenCostDetectability
API (incl. internal XHR/JSON endpoints)Any documented API, or the page fetches its own JSON you can call directlyLowestLowest — looks like normal traffic
sitemap.xml + JSON-LD/sitemap.xml lists the URLs; <script type="application/ld+json"> carries the records (JSON-LD is Google's preferred structured format, so it is everywhere)Low — one HTTP GET, parse JSONLow
HTML selectorsData is in server-rendered HTML, no JSON-LD, no XHR JSONMedium — selectors drift on redesignMedium
Headless browserThe data only exists after JS executes (SPA, lazy-load, infinite scroll) and there is no callable XHR endpointHighest — CPU, RAM, time, easiest to fingerprintHighest

Before you reach for a browser, open DevTools → Network and look for the XHR/fetch that already returns JSON. Calling that endpoint directly is faster, stabler, and less detectable than rendering the whole page to read what the page itself fetched.

Tool picker

Site profileToolVersion (2026)The one reason
Static HTML, no fingerprint wallhttpx + selectolax (or BeautifulSoup)currentFast, no browser; selectolax parses far quicker than lxml
Static HTML but TLS/JA3 fingerprint blocks youcurl_cffi (profile chrome131)currentImpersonates a real browser's TLS/JA3/HTTP2 fingerprint, not just the User-Agent
JS-rendered SPAPlaywright1.60.0 (1.59 shipped 2026-04-01)First-class async, auto-wait, the maintained headless standard
Scalable, resilient, recurring crawlerCrawlee (JS or Python)actively maintained 2026Wraps Playwright + proxy rotation + browserforge fingerprints + a disk-persisted RequestQueue that resumes after a crash
Pure-Python static, legacy codebaseScrapycurrentBattle-tested for static targets — but its Twisted core lags the asyncio ecosystem; pick Crawlee for new work

Default new builds to Crawlee when the job is recurring or must not break; reach for plain httpx/curl_cffi only when the target is static and one-shot.

Robust selectors

Selectors break on redesign because they ride on layout, not meaning. Anchor on what is semantically stable — data-* attributes, ARIA roles, microdata, visible text — never on nth-child chains or generated CSS class hashes (.css-1a2b3c), which change on every build.

html
<!-- the page you are scraping -->
<article data-testid="listing-card">
  <h2 class="css-1a2b3c">Acme Drill 9000</h2>
  <span data-price="129.00">€129,00</span>
</article>
python
# Bad — rides on layout and a build-generated hash; dies on the next deploy
title = page.query_selector("div:nth-child(3) > article > .css-1a2b3c").inner_text()

# Good — anchor on stable semantics, with a fallback chain, and fail loud
def text_or_raise(card, selectors, field):
    for sel in selectors:           # try each selector in priority order
        el = card.query_selector(sel)
        if el and el.inner_text().strip():
            return el.inner_text().strip()
    raise LookupError(f"required field {field!r} not found via {selectors}")

card  = page.query_selector('[data-testid="listing-card"]')
title = text_or_raise(card, ["h2", '[itemprop="name"]'], "title")
price = text_or_raise(card, ["[data-price]", "span:has-text('€')"], "price")

Two rules baked into that snippet:

  • Fallback selector chains. Give every required field 2-3 ordered selectors. A single redesign rarely breaks all of them at once, so the scraper degrades instead of dying.
  • Fail loud on a missing required field. Raise — never silently write null. A silent null is a corrupted dataset you discover three months later. A raised error is a fix you make today.
Show full SKILL.md (686 more words)Show less

Politeness and anti-block

Concrete numbers beat "be respectful." Depth — fingerprint profiles, header sets, proxy taxonomy — is in references/anti-bot.md.

  • Pace. Start at 1 request per 1-3 seconds per host, concurrency capped at 2-5 per host. Tune down if you see 429s, never up to chase speed.
  • Backoff with jitter. On a transient error, exponential backoff plus random jitter — delay = base * 2**attempt + random(0, base). Without jitter, every worker retries in lockstep and you self-DDoS the host.
  • Honor 429 and Retry-After. A 429 with Retry-After: 30 means wait 30 seconds, not retry immediately. Ignoring it is the fastest route to an IP ban.
  • Conditional GET. Send If-Modified-Since / If-None-Match (ETag) on re-crawls. A 304 Not Modified costs nothing and tells you the page is unchanged — cheaper for you, lighter on the host.
  • Realistic headers + TLS fingerprint. A bare python-requests User-Agent is an instant tell. Send a full, current browser header set; when JA3/TLS fingerprinting blocks you, switch to curl_cffi (chrome131) — see the references.
  • Proxies only when justified. Residential/mobile proxies are for geo-gating and IP-rate distribution on a target that permits the scrape — never to evade a block you were explicitly shown. Budget for them on large recurring jobs; do not reach for them to defeat a hard block (that is the legal line you do not cross).

Resilience beats speed. A scraper that runs 50% slower but never breaks is infinitely more valuable than a fast one that dies weekly. Pace for survival, not throughput.

Resilience

A recurring crawler must survive crashes, redeploys, and the target's schema drift. Patterns and copy-paste starters are in references/frameworks.md.

  • Resumable queue. Use a disk-persisted request queue (Crawlee's RequestQueue) so a crash resumes from where it stopped, not from zero. Re-crawling 10k pages because the box rebooted is wasted budget and extra load on the host.
  • Idempotent writes. Upsert keyed on a stable id (the source URL or a record id), never blind append — so a re-run repairs rows instead of duplicating them.
  • Checkpoint every N records. Flush progress periodically so an interrupt loses minutes, not hours.
  • Change detection by hash. Store a content hash per record; only re-process when it changes. This pairs with conditional GET to keep recurring crawls cheap.
  • Monitor and alert. Alert on a spike in missing-field errors (schema drift — the site redesigned) and on a spike in 403/429 (you are being blocked). Both are silent-failure modes that rot a dataset until someone checks.

Anti-patterns

Anti-patternWhy it bitesDo instead
Scraping behind a login, then calling it "public data"Accepting ToS at login is the breach-of-contract hook (Meta v. Bright Data)Stay logged-out on public pages; never bypass auth
Not even reading robots.txt / ai.txtYou lose free intelligence on where you will get blocked, and any good-faith story laterParse both; honor by default, override deliberately and write down why
No delay, unbounded concurrencyHammers the host → IP ban, possible CFAA-style exposure1 req / 1-3s, cap 2-5 concurrent per host
Selectors on nth-child / .css-1a2b3c hashesBreak on the next deploy; silent data lossAnchor on data-* / semantic / text, with fallbacks
Silently writing null on a missing fieldCorrupts the dataset; discovered months laterRaise on a missing required field — fail loud
Solving a CAPTCHA / bypassing a hard blockCircumventing a shown control (Reddit v. Perplexity) — the worst legal postureStop. That control is a "no."
Scraping personal data with no lawful basisGDPR fines into tens of millions EUREstablish basis + minimize + filter special categories (../gdpr-privacy/SKILL.md)
Storing everything "just in case"Defeats minimization; expands breach blast radiusKeep only fields the purpose needs; set retention
Launching a browser when JSON-LD was right thereSlowest, most detectable, most expensive pathCheck XHR/JSON-LD/sitemap first; browser is last
No resume — a crash restarts from zeroWastes budget, doubles load on the hostDisk-persisted resumable queue
Hardcoding one User-Agent foreverStale UA is an easy bot tellCurrent full header set; rotate when justified
Retrying with no backoffLockstep retries self-DDoS the hostExponential backoff + jitter; honor Retry-After

When in doubt about whether a scrape is defensible, the answer is the gate. Run it, write down the outcome, and only then send a request.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/data-scraper of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/anti-bot.md
  • references/frameworks.md
  • references/legal-compliance.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Data Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Scraper this skillericrisco/rsc-harness156—~3.3kAutomated safety check: PassMIT
Cheerio ParsingKilo-Org/kilo-marketplace1891 repos~2.2kAutomated safety check: PassApache-2.0
Pp Context Devmvanhorn/printing-press-library2.1k—~4.3kAutomated safety check: NotesApache-2.0
Steel Browseraiskillstore/marketplace430—~2kAutomated safety check: PassNone
Tmuxtrpc-group/trpc-agent-go1.8k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6961 repos~3.9kAutomated safety check: PassMIT

Similar skills

  • Cheerio Parsing

    Kilo-Org/kilo-marketplace

    Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.

    189 GitHub starsUsed in 1 repo~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Pp Context Dev

    mvanhorn/printing-press-library

    Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library.

    2.1k GitHub stars~4.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check: notes
  • Steel Browser

    aiskillstore/marketplace

    Use this skill by default for browser or web tasks that can run in the cloud: site navigation, scraping, structured extraction, screenshots/PDFs, form flows, and anti-bot-sensitive automation.

    430 GitHub stars~2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    696 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Questions about Data Scraper

What does Data Scraper do?

A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…. Data Scraper is an agent skill from ericrisco/rsc-harness. Use when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable selectors, pacing, resilience.

When should I use Data Scraper?

Data Scraper fits situations like: data lives on a website with no usable API — listings; public records — and the scrape must stay legal and not get blocked: legal gate; extraction path; durable selectors.

How do I install Data Scraper in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill data-scraper -a claude-code`. Or copy the skill folder (skills/data-scraper in ericrisco/rsc-harness) into .claude/skills/data-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Data Scraper in Codex?

Run `npx skills add ericrisco/rsc-harness --skill data-scraper -a codex`. Or copy the skill folder (skills/data-scraper in ericrisco/rsc-harness) into .agents/skills/data-scraper in your project. Codex loads it when a task matches its description.

Can I use Data Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill data-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scraper, .gemini/skills/data-scraper, .github/skills/data-scraper and .opencode/skills/data-scraper in your project.

What does Data Scraper need to run?

Going by SKILL.md and its folder, Data Scraper needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Data Scraper access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Scraper use?

Data Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Scraper use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Data Scraper?

Skills that share tags, products or a category with Data Scraper: Cheerio Parsing (Kilo-Org/kilo-marketplace, 189 stars), Pp Context Dev (mvanhorn/printing-press-library, 2.1k stars), Steel Browser (aiskillstore/marketplace, 430 stars) and Tmux (trpc-group/trpc-agent-go, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Scraper?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.