Agent skill

Cashclaw Data Scraper

by ertugrulakben in ertugrulakben/cashclaw

Extracts structured data from websites and APIs, delivering clean datasets in multiple formats.

MITAuto-check passedData & Analytics

Install Cashclaw Data Scraper

skills CLI
$ npx skills add ertugrulakben/cashclaw --skill cashclaw-data-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ertugrulakben/cashclaw cashclaw-data-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ertugrulakben/cashclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cashclaw-data-scraper .claude/skills/cashclaw-data-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cashclaw-data-scraper
GitHub stars
303
Token cost
~3k tokens
SKILL.md length
602 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Extracts structured data from websites and APIs, delivering clean datasets in multiple formats.

  • Works in 6 steps: Source Identification → Schema Definition → Data Extraction → …
  • Tasks that involve Web scraping
  • SKILL.md covers Pricing Tiers, Data Extraction Workflow, Quality Checklist and Deliverable Format, plus 3 more sections
  • Calls curl, node and jq; reaches acme.com and beta.io

What it does

Cashclaw Data Scraper is an agent skill from ertugrulakben/cashclaw. Extracts structured data from websites and APIs, delivering clean datasets in multiple formats. Handles pagination, deduplication, and data enrichment for reliable business intelligence.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping, Data cleaning and Schema markup. The repository describes itself as: The Agent Economy Layer — agents earn, agents spend, Guard protects. 13 skills, runtime cost cap, recursive kill, tool firewall. 50+ HYRVE API endpoints, job polling daemon, MPP…. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve Data cleaning
  • Tasks that involve Schema markup

Example prompts

  • “Use the cashclaw-data-scraper skill to extract structured data from websites and APIs, delivering clean datasets in multiple formats”
  • “/cashclaw-data-scraper”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Source Identification
  2. Schema Definition
  3. Data Extraction
  4. Data Cleaning and Deduplication
  5. Data Enrichment (Pro Tier)
  6. Export and Delivery

What it can do on your machine

Read from SKILL.md and the folder at commit ff30cb3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • node
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • acme.com
    • beta.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cashclaw Data Scraper loads about 3k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 602 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ertugrulakben/cashclaw at commit ff30cb3, republished under its MIT licence (© ertugrulakben). 602 words, ~2,987 tokens.

Download SKILL.mdSave it as .claude/skills/cashclaw-data-scraper/SKILL.md (or your agent's skills folder).
name
cashclaw-data-scraper
description
Extracts structured data from websites and APIs, delivering clean datasets in multiple formats. Handles pagination, deduplication, and data enrichment for reliable business intelligence.

CashClaw Data Scraper

You extract structured data from websites and APIs that clients need for business decisions. Every dataset must be clean, deduplicated, and delivered in the requested format. Raw unprocessed dumps are not deliverables. Quality and accuracy matter more than volume.

Pricing Tiers

TierScopePriceDelivery
BasicSingle source, up to 50 records$93 hours
StandardMultiple sources, up to 200 records, dedup$1912 hours
ProMultiple sources, up to 500 records + enrichment$2524 hours

Data Extraction Workflow

Step 1: Source Identification

When you receive a scraping request, extract or ask for:

  1. Target URL(s) - Specific pages, search results, directories, or API endpoints.
  2. Data Fields - Exactly what data points are needed (name, email, price, etc.).
  3. Record Count - How many records does the client need?
  4. Output Format - CSV, JSON, or both.
  5. Filters - Any criteria to include/exclude records (geography, category, price range).
  6. Freshness - Does the data need to be current, or is historical data acceptable?
  7. Update Frequency - One-time extraction or recurring?
  8. Use Case - What will the data be used for? (This affects what is ethical to collect.)

If the client says "scrape everything from this site," push back and ask for specific fields and record limits. Unbounded scraping is irresponsible.

Step 2: Schema Definition

Before extracting any data, define the output schema:

json
{
  "$schema": "extraction-schema-v1",
  "source": "{source_url}",
  "description": "{what this dataset contains}",
  "fields": [
    {
      "name": "company_name",
      "type": "string",
      "required": true,
      "description": "Legal company name"
    },
    {
      "name": "website",
      "type": "url",
      "required": true,
      "description": "Company website URL"
    },
    {
      "name": "industry",
      "type": "string",
      "required": false,
      "description": "Primary industry category"
    },
    {
      "name": "employee_count",
      "type": "integer",
      "required": false,
      "description": "Approximate employee count"
    },
    {
      "name": "location",
      "type": "string",
      "required": false,
      "description": "Headquarters city, state/country"
    }
  ],
  "dedup_key": "website",
  "sort_by": "company_name",
  "filters": {
    "industry": "{filter_value}",
    "min_employees": 10
  }
}

Share this schema with the client for approval before extraction begins.

Step 3: Data Extraction

Use the appropriate extraction method based on the source:

Method A: API-Based Extraction (preferred)

bash
# If the source has a public API
curl -s "https://api.example.com/v1/companies?industry=saas&limit=50" \
  -H "Accept: application/json" | jq '.data[]' > raw-data.json

Method B: HTML Scraping

bash
# Fetch the page
curl -sL "https://example.com/directory?page=1" -o page.html

# Parse with node script
node scripts/scraper.js --url "https://example.com/directory" --pages 5 --output raw-data.json

Method C: Structured Data Extraction

bash
# Extract JSON-LD, microdata, or Open Graph from pages
node scripts/extract-structured.js --url "https://example.com" --format jsonld

Pagination Handling:

When the data spans multiple pages:

  1. Identify the pagination pattern (page numbers, offset, cursor, next URL).
  2. Calculate total pages needed: ceil(target_records / records_per_page).
  3. Add a 1-2 second delay between requests to respect server resources.
  4. Handle pagination edge cases: empty pages, duplicate last-page entries.
yaml
Pagination Config:
  Pattern: "{query_param | path | cursor | link_header}"
  Base URL: "{url}"
  Page Param: "page={n}"
  Records Per Page: 20
  Total Pages Needed: 3
  Delay Between Requests: 1500ms
  Stop Condition: "empty results OR target count reached"
Step 4: Data Cleaning and Deduplication

Apply these cleaning steps to every dataset:

yaml
Cleaning Pipeline:
  1. Remove Duplicates:
     - Deduplicate on primary key (e.g., website domain)
     - If two records share the same key, keep the more complete one

  2. Normalize Fields:
     - URLs: Add https:// if missing, remove trailing slashes
     - Phone: Standardize to E.164 format (+1XXXXXXXXXX)
     - Email: Lowercase, trim whitespace
     - Company Names: Trim, normalize casing (Title Case)
     - Locations: Standardize to "City, State, Country" format

  3. Validate Data Types:
     - URLs: Must start with http:// or https://
     - Emails: Must match RFC 5322 pattern
     - Numbers: Must be numeric (remove currency symbols, commas)
     - Dates: Normalize to ISO 8601

  4. Handle Missing Data:
     - Required fields missing: Flag record for review or discard
     - Optional fields missing: Set to null, not empty string
     - Never fabricate data to fill gaps

  5. Quality Score:
     - Calculate completeness percentage per record
     - Flag records below 60% completeness for review
Step 5: Data Enrichment (Pro Tier)

For Pro tier, enrich the base dataset with additional data points:

yaml
Enrichment Sources:
  Company Data:
    - Employee count from LinkedIn company page
    - Industry classification from website metadata
    - Tech stack from BuiltWith or Wappalyzer signals
    - Social media profiles from website footer links

  Contact Data:
    - Email pattern detection (first@, first.last@, firstl@)
    - LinkedIn profile URLs from company team page
    - Phone from website contact page

  Business Signals:
    - Recent funding (Crunchbase, press releases)
    - Job openings count (careers page, job boards)
    - Website traffic estimate (if observable)
    - Social media activity level

Mark all enriched fields with their source and confidence level:

json
{
  "company_name": "Acme Corp",
  "website": "https://acme.com",
  "enriched": {
    "employee_count": {
      "value": 85,
      "source": "linkedin",
      "confidence": "high",
      "date": "2026-03-15"
    },
    "tech_stack": {
      "value": ["React", "Node.js", "AWS"],
      "source": "website_analysis",
      "confidence": "medium",
      "date": "2026-03-15"
    }
  }
}
Show full SKILL.md (254 more words)Show less
Step 6: Export and Delivery

Package the data in the requested format(s):

CSV Output:

csv
company_name,website,industry,employee_count,location,email,phone,score
"Acme Corp","https://acme.com","SaaS",85,"Austin, TX","info@acme.com","+15550123",92
"Beta Inc","https://beta.io","Fintech",42,"New York, NY","hello@beta.io","+15550456",87

CSV rules:

  • UTF-8 encoding with BOM for Excel compatibility.
  • Quote all string fields.
  • Use comma delimiter (not semicolon or tab).
  • Header row required.
  • No trailing commas.

JSON Output:

json
{
  "metadata": {
    "source": "{source_url}",
    "extracted_at": "{ISO8601}",
    "total_records": 50,
    "schema_version": "1.0",
    "completeness_avg": 87,
    "dedup_applied": true
  },
  "records": [
    {
      "company_name": "Acme Corp",
      "website": "https://acme.com",
      "industry": "SaaS",
      "employee_count": 85,
      "location": "Austin, TX",
      "quality_score": 92
    }
  ]
}

Quality Checklist

Before delivering, verify:

[ ] Record count matches the tier (50 / 200 / 500)
[ ] No duplicate records (verified on dedup key)
[ ] All required fields are populated
[ ] URLs are valid and accessible
[ ] Email addresses pass format validation
[ ] Phone numbers are in consistent format
[ ] No obviously stale data (defunct companies, dead links)
[ ] CSV opens correctly in Excel/Google Sheets
[ ] JSON is valid (passes a linter)
[ ] Completeness score average is above 75%
[ ] Enrichment sources are documented (Pro tier)
[ ] Extraction report includes methodology
[ ] No personally identifiable information beyond business context
[ ] Data is sorted according to schema definition
[ ] Character encoding is UTF-8 throughout

Deliverable Format

Every data extraction delivery includes:

deliverables/
  data-{source}-{date}.csv              - Clean dataset in CSV
  data-{source}-{date}.json             - Clean dataset in JSON
  extraction-report.md                   - Methodology, stats, quality notes
extraction-report.md Format
markdown
# Data Extraction Report

**Source:** {source_url}
**Date:** {date}
**Tier:** {Basic|Standard|Pro}

## Summary
- Records Requested: {count}
- Records Delivered: {count}
- Completeness Average: {percent}%
- Duplicates Removed: {count}

## Schema
| Field | Type | Required | Population Rate |
|-------|------|----------|-----------------|
| company_name | string | yes | 100% |
| website | url | yes | 100% |
| industry | string | no | 85% |
| employee_count | integer | no | 72% |

## Methodology
- Sources used: {list}
- Pages scraped: {count}
- Extraction method: {API / HTML parsing / structured data}
- Deduplication key: {field}

## Data Quality Notes
- {Any issues encountered}
- {Fields with low population rates and why}
- {Recommendations for improving data quality}

## Ethical Compliance
- robots.txt respected: {yes/no}
- Rate limiting applied: {delay between requests}
- Terms of service reviewed: {compliant/concerns noted}

Ethical Guidelines

These rules are non-negotiable:

  1. Respect robots.txt - Check and honor robots.txt directives before scraping.
  2. Rate limiting - Minimum 1 second delay between requests. Never DDoS a site.
  3. No authentication bypass - Do not circumvent login walls, CAPTCHAs, or paywalls.
  4. No personal data - Do not scrape personal social media profiles, home addresses, or private information.
  5. Business context only - Collect only business-relevant data (company info, public business contacts).
  6. Terms of service - Review the site's ToS. Flag concerns to the client if scraping may violate them.
  7. No resale of scraped data - Data is for the client's internal use only unless otherwise cleared.
  8. Attribution - Note the source of every data point in the extraction report.

Quality Standards

  • Every record must have all required fields populated.
  • Completeness average must be 75% or higher.
  • Zero duplicates in the delivered dataset.
  • Data must be current -- no records older than 90 days unless historical data was requested.
  • If the target record count cannot be met from the specified source, deliver what is available and explain the shortfall. Never pad with fabricated records.
  • Pro tier enrichment must add at least 3 new data points per record on average.

Example Commands

bash
# Basic extraction from a single source
cashclaw scrape --url "https://directory.example.com/companies" --fields "name,website,industry" --limit 50 --output data.csv

# Standard multi-source extraction
cashclaw scrape --urls "source1.com/list,source2.com/directory" --fields "name,website,email,phone" --limit 200 --dedup website --output data.json

# Pro extraction with enrichment
cashclaw scrape --url "https://directory.example.com" --fields "name,website,industry,size" --limit 500 --enrich --output data.csv data.json

# Validate an existing dataset
cashclaw scrape validate --input data.csv --schema schema.json --report quality-report.md

© ertugrulakben, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cashclaw-data-scraper of ertugrulakben/cashclaw.

Open the folder on GitHubat commit ff30cb3

Compare with similar skills

Cashclaw Data Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cashclaw Data Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cashclaw Data Scraper this skillertugrulakben/cashclaw303—~3kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT
Meowhub Browserzhaojiaqi/MeowHub111—~1.6kAutomated safety check: PassGPL-3.0
Authoritative Data Harvesteryushui2022/MathModel-Skill4541 repos~1.1kAutomated safety check: PassMIT
Chatgpt SearchSeifBenayed/cloclo114—~1.7kAutomated safety check: NotesMIT
Firecrawl Agentfirecrawl/skills117—~1.2kAutomated safety check: PassISC

Similar skills

  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Meowhub Browser

    zhaojiaqi/MeowHub

    Browse the web using Browserless.io cloud browser service. An agent skill from zhaojiaqi/MeowHub.

    111 GitHub stars~1.6k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    454 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Chatgpt Search

    SeifBenayed/cloclo

    Search ChatGPT and extract the full response + hydration JSON that powers the UI.

    114 GitHub stars~1.7k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check: notes
  • Firecrawl Agent

    firecrawl/skills

    Autonomously navigate websites and extract structured data across pages.

    117 GitHub stars~1.2k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit agent readiness by script: AI-crawler rules, product schema, no-JS HTML, feeds.

    862 GitHub starsUsed in 1 repo~3.7k tokens
    Data & AnalyticsAuto-check passed

More from ertugrulakben/cashclaw

All 13 skills in this repo
  • Cashclaw Guard

    ertugrulakben/cashclaw

    Runtime protection layer for AI agents. An agent skill from ertugrulakben/cashclaw.

    303 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed
  • Cashclaw Invoicer

    ertugrulakben/cashclaw

    Handles invoice creation, payment link generation, payment status tracking, and automated reminders via Stripe API.

    303 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Cashclaw Lead Generator

    ertugrulakben/cashclaw

    Generates qualified B2B leads through systematic research, data collection, and scoring.

    303 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Cashclaw SEO Auditor

    ertugrulakben/cashclaw

    Performs comprehensive SEO audits on websites covering technical SEO, on-page optimization, off-page signals, and performance metrics.

    303 GitHub stars~1.7k tokensUpdated 4 days ago
    Auto-check passed
  • Cashclaw Competitor Analyzer

    ertugrulakben/cashclaw

    Performs competitor research and generates detailed analysis reports with market positioning insights.

    303 GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed
  • Cashclaw Content Writer

    ertugrulakben/cashclaw

    Writes professional blog posts, social media content, and email newsletters optimized for SEO and engagement.

    303 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed

Questions about Cashclaw Data Scraper

What does Cashclaw Data Scraper do?

Extracts structured data from websites and APIs, delivering clean datasets in multiple formats. Cashclaw Data Scraper is an agent skill from ertugrulakben/cashclaw. Extracts structured data from websites and APIs, delivering clean datasets in multiple formats.

When should I use Cashclaw Data Scraper?

Cashclaw Data Scraper fits situations like: tasks that involve Web scraping; tasks that involve Data cleaning; tasks that involve Schema markup.

How do I install Cashclaw Data Scraper in Claude Code?

Run `npx skills add ertugrulakben/cashclaw --skill cashclaw-data-scraper -a claude-code`. Or copy the skill folder (skills/cashclaw-data-scraper in ertugrulakben/cashclaw) into .claude/skills/cashclaw-data-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Cashclaw Data Scraper in Codex?

Run `npx skills add ertugrulakben/cashclaw --skill cashclaw-data-scraper -a codex`. Or copy the skill folder (skills/cashclaw-data-scraper in ertugrulakben/cashclaw) into .agents/skills/cashclaw-data-scraper in your project. Codex loads it when a task matches its description.

Can I use Cashclaw Data Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ertugrulakben/cashclaw --skill cashclaw-data-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cashclaw-data-scraper, .gemini/skills/cashclaw-data-scraper, .github/skills/cashclaw-data-scraper and .opencode/skills/cashclaw-data-scraper in your project.

What does Cashclaw Data Scraper need to run?

Going by SKILL.md and its folder, Cashclaw Data Scraper needs the command-line tools its instructions call (curl, node and jq). Our summary lists: Node.js.

Does Cashclaw Data Scraper access the network?

SKILL.md names 2 domains. In commands or code: acme.com and beta.io; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Cashclaw Data Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cashclaw Data Scraper use?

Cashclaw Data Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cashclaw Data Scraper use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cashclaw Data Scraper?

Skills that share tags, products or a category with Cashclaw Data Scraper: Ax (yusukebe/ax, 719 stars), Meowhub Browser (zhaojiaqi/MeowHub, 111 stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 454 stars) and Chatgpt Search (SeifBenayed/cloclo, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cashclaw Data Scraper?

ertugrulakben (a GitHub user) maintains it in ertugrulakben/cashclaw, which has 303 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 6, 2026.

Source: ertugrulakben/cashclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.