Cheerio Parsing
Kilo-Org/kilo-marketplace
Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.
A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness data-scraper --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-scraper .claude/skills/data-scraper && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .claude/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraperType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness data-scraper --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data-scraper .agents/skills/data-scraper && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .agents/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness data-scraper --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data-scraper .cursor/skills/data-scraper && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .cursor/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/data-scraper--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness data-scraper --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data-scraper .gemini/skills/data-scraper && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .gemini/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness data-scraperInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data-scraper .github/skills/data-scraper && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .github/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill data-scraper -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness data-scraper --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data-scraper .opencode/skills/data-scraper && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-scraper" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-scraper into .opencode/skills/data-scraper/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-scraper", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-scraperA skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…
Data Scraper is an agent skill from ericrisco/rsc-harness. Use when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable selectors, pacing, resilience. NOT parsing bytes you already hold into fields (that is structured-extraction), NOT a documented API or key (that is api-connector-builder).
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/anti-bot.md`).
It sits in Data & Analytics, covering Web scraping and Document parsing. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Scraper loads about 3.3k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 1,708 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,708 words, ~3,310 tokens.
.claude/skills/data-scraper/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.You get bytes off websites you do not control — legally, and without getting blocked. You pick the cheapest extraction path that works, build selectors that survive a redesign, pace requests so the host neither bans nor sues you, and you write down the legal basis before the first request goes out.
One rule above all the others: scraping is the fallback, not the default. It is what you reach for only when no API serves the data. If the site has a documented API or you hold a key, stop — that is ../api-connector-builder/SKILL.md. And once you have the bytes, parsing them into fields is ../structured-extraction/SKILL.md, normalizing the rows is ../data-cleaning/SKILL.md. This skill ends the moment you hold the bytes.
In 2025-2026 scrapers do not fail on parsing. They fail on a terms-of-service breach, on GDPR exposure, or on being blocked after hammering a host. So the work runs in this order: legal gate → extraction path → tool → selectors → politeness → resilience. Skipping the gate is how you end up in Meta v. Bright Data.
Walk every item. Each ends in proceed, proceed-narrowed, or stop. One item at stop means the whole scrape stops until you resolve it. Depth and the case law are in references/legal-compliance.md.
The gates below are the ones with real legal exposure — contract, data-protection law, and anti-circumvention. robots.txt is not one of them: it is a voluntary convention with no statutory force, so it informs the decision and never blocks it on its own.
../gdpr-privacy/SKILL.md), or stop.Disallow tells you the host would rather you did not, and Crawl-delay tells you its tolerance; both are useful intelligence about where you are likely to get blocked. Weigh it: honoring robots is the low-friction default and the cleanest evidence of good faith, but scraping your own site, one you have permission for, or public pages for research is legitimate whether or not robots allows it. → advisory: note the decision and move on.hiQ v. LinkedIn established that scraping public data is not automatically a CFAA violation — but that is the floor, not a license. Public + logged-out + no-personal-data is the defensible quadrant, and honoring robots keeps it tidy without being what makes it lawful. Anything outside it, document why and get a human to sign off.
Walk down only when the rung above is unavailable. ~94% of modern sites are client-side rendered, yet most still ship machine-readable structured data — parse that before you launch a browser.
| Path | Use when | Cost | Detectability |
|---|---|---|---|
| API (incl. internal XHR/JSON endpoints) | Any documented API, or the page fetches its own JSON you can call directly | Lowest | Lowest — looks like normal traffic |
| sitemap.xml + JSON-LD | /sitemap.xml lists the URLs; <script type="application/ld+json"> carries the records (JSON-LD is Google's preferred structured format, so it is everywhere) | Low — one HTTP GET, parse JSON | Low |
| HTML selectors | Data is in server-rendered HTML, no JSON-LD, no XHR JSON | Medium — selectors drift on redesign | Medium |
| Headless browser | The data only exists after JS executes (SPA, lazy-load, infinite scroll) and there is no callable XHR endpoint | Highest — CPU, RAM, time, easiest to fingerprint | Highest |
Before you reach for a browser, open DevTools → Network and look for the XHR/fetch that already returns JSON. Calling that endpoint directly is faster, stabler, and less detectable than rendering the whole page to read what the page itself fetched.
| Site profile | Tool | Version (2026) | The one reason |
|---|---|---|---|
| Static HTML, no fingerprint wall | httpx + selectolax (or BeautifulSoup) | current | Fast, no browser; selectolax parses far quicker than lxml |
| Static HTML but TLS/JA3 fingerprint blocks you | curl_cffi (profile chrome131) | current | Impersonates a real browser's TLS/JA3/HTTP2 fingerprint, not just the User-Agent |
| JS-rendered SPA | Playwright | 1.60.0 (1.59 shipped 2026-04-01) | First-class async, auto-wait, the maintained headless standard |
| Scalable, resilient, recurring crawler | Crawlee (JS or Python) | actively maintained 2026 | Wraps Playwright + proxy rotation + browserforge fingerprints + a disk-persisted RequestQueue that resumes after a crash |
| Pure-Python static, legacy codebase | Scrapy | current | Battle-tested for static targets — but its Twisted core lags the asyncio ecosystem; pick Crawlee for new work |
Default new builds to Crawlee when the job is recurring or must not break; reach for plain httpx/curl_cffi only when the target is static and one-shot.
Selectors break on redesign because they ride on layout, not meaning. Anchor on what is semantically stable — data-* attributes, ARIA roles, microdata, visible text — never on nth-child chains or generated CSS class hashes (.css-1a2b3c), which change on every build.
<!-- the page you are scraping -->
<article data-testid="listing-card">
<h2 class="css-1a2b3c">Acme Drill 9000</h2>
<span data-price="129.00">€129,00</span>
</article># Bad — rides on layout and a build-generated hash; dies on the next deploy
title = page.query_selector("div:nth-child(3) > article > .css-1a2b3c").inner_text()
# Good — anchor on stable semantics, with a fallback chain, and fail loud
def text_or_raise(card, selectors, field):
for sel in selectors: # try each selector in priority order
el = card.query_selector(sel)
if el and el.inner_text().strip():
return el.inner_text().strip()
raise LookupError(f"required field {field!r} not found via {selectors}")
card = page.query_selector('[data-testid="listing-card"]')
title = text_or_raise(card, ["h2", '[itemprop="name"]'], "title")
price = text_or_raise(card, ["[data-price]", "span:has-text('€')"], "price")Two rules baked into that snippet:
null. A silent null is a corrupted dataset you discover three months later. A raised error is a fix you make today.Concrete numbers beat "be respectful." Depth — fingerprint profiles, header sets, proxy taxonomy — is in references/anti-bot.md.
delay = base * 2**attempt + random(0, base). Without jitter, every worker retries in lockstep and you self-DDoS the host.Retry-After. A 429 with Retry-After: 30 means wait 30 seconds, not retry immediately. Ignoring it is the fastest route to an IP ban.If-Modified-Since / If-None-Match (ETag) on re-crawls. A 304 Not Modified costs nothing and tells you the page is unchanged — cheaper for you, lighter on the host.python-requests User-Agent is an instant tell. Send a full, current browser header set; when JA3/TLS fingerprinting blocks you, switch to curl_cffi (chrome131) — see the references.Resilience beats speed. A scraper that runs 50% slower but never breaks is infinitely more valuable than a fast one that dies weekly. Pace for survival, not throughput.
A recurring crawler must survive crashes, redeploys, and the target's schema drift. Patterns and copy-paste starters are in references/frameworks.md.
RequestQueue) so a crash resumes from where it stopped, not from zero. Re-crawling 10k pages because the box rebooted is wasted budget and extra load on the host.| Anti-pattern | Why it bites | Do instead |
|---|---|---|
| Scraping behind a login, then calling it "public data" | Accepting ToS at login is the breach-of-contract hook (Meta v. Bright Data) | Stay logged-out on public pages; never bypass auth |
| Not even reading robots.txt / ai.txt | You lose free intelligence on where you will get blocked, and any good-faith story later | Parse both; honor by default, override deliberately and write down why |
| No delay, unbounded concurrency | Hammers the host → IP ban, possible CFAA-style exposure | 1 req / 1-3s, cap 2-5 concurrent per host |
Selectors on nth-child / .css-1a2b3c hashes | Break on the next deploy; silent data loss | Anchor on data-* / semantic / text, with fallbacks |
Silently writing null on a missing field | Corrupts the dataset; discovered months later | Raise on a missing required field — fail loud |
| Solving a CAPTCHA / bypassing a hard block | Circumventing a shown control (Reddit v. Perplexity) — the worst legal posture | Stop. That control is a "no." |
| Scraping personal data with no lawful basis | GDPR fines into tens of millions EUR | Establish basis + minimize + filter special categories (../gdpr-privacy/SKILL.md) |
| Storing everything "just in case" | Defeats minimization; expands breach blast radius | Keep only fields the purpose needs; set retention |
| Launching a browser when JSON-LD was right there | Slowest, most detectable, most expensive path | Check XHR/JSON-LD/sitemap first; browser is last |
| No resume — a crash restarts from zero | Wastes budget, doubles load on the host | Disk-persisted resumable queue |
| Hardcoding one User-Agent forever | Stale UA is an easy bot tell | Current full header set; rotate when justified |
| Retrying with no backoff | Lockstep retries self-DDoS the host | Exponential backoff + jitter; honor Retry-After |
When in doubt about whether a scrape is defensible, the answer is the gate. Run it, write down the outcome, and only then send a request.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in skills/data-scraper of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Data Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Scraper this skillericrisco/rsc-harness | 156 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Cheerio ParsingKilo-Org/kilo-marketplace | 189 | 1 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Pp Context Devmvanhorn/printing-press-library | 2.1k | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | |
| Steel Browseraiskillstore/marketplace | 430 | — | ~2k | Automated safety check: Pass | None | |
| Tmuxtrpc-group/trpc-agent-go | 1.8k | 23 repos | ~868 | Automated safety check: Pass | Apache-2.0 | |
| Ketch1broseidon/ketch | 696 | 1 repos | ~3.9k | Automated safety check: Pass | MIT |
Kilo-Org/kilo-marketplace
Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.
mvanhorn/printing-press-library
Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library.
aiskillstore/marketplace
Use this skill by default for browser or web tasks that can run in the cloud: site navigation, scraping, structured extraction, screenshots/PDFs, form flows, and anti-bot-sensitive automation.
trpc-group/trpc-agent-go
Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.
1broseidon/ketch
Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Categories
A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…. Data Scraper is an agent skill from ericrisco/rsc-harness. Use when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable selectors, pacing, resilience.
Data Scraper fits situations like: data lives on a website with no usable API — listings; public records — and the scrape must stay legal and not get blocked: legal gate; extraction path; durable selectors.
Run `npx skills add ericrisco/rsc-harness --skill data-scraper -a claude-code`. Or copy the skill folder (skills/data-scraper in ericrisco/rsc-harness) into .claude/skills/data-scraper in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill data-scraper -a codex`. Or copy the skill folder (skills/data-scraper in ericrisco/rsc-harness) into .agents/skills/data-scraper in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill data-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scraper, .gemini/skills/data-scraper, .github/skills/data-scraper and .opencode/skills/data-scraper in your project.
Going by SKILL.md and its folder, Data Scraper needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Scraper: Cheerio Parsing (Kilo-Org/kilo-marketplace, 189 stars), Pp Context Dev (mvanhorn/printing-press-library, 2.1k stars), Steel Browser (aiskillstore/marketplace, 430 stars) and Tmux (trpc-group/trpc-agent-go, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.