Agent skill

Universal Scraping Architect

by alirezarezvani in alirezarezvani/claude-skills

A skill your agent uses for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

MITAuto-check passedData & Analytics

Install Universal Scraping Architect

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill universal-scraping-architect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills universal-scraping-architect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/universal-scraping-architect/skills/universal-scraping-architect .claude/skills/universal-scraping-architect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
universal-scraping-architect
GitHub stars
28k
Token cost
~1.1k tokens
SKILL.md length
558 words
Files
8 (incl. scripts, references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

  • Works in 5 steps: Route the Approach: Explicitly state… → Track Budgets: Estimate Firecrawl API… → Extract Safely: Implement checkpointing… → …
  • Document extraction
  • SKILL.md covers Before Starting, How This Skill Works, The Extraction Pipeline and Proactive Triggers, plus 3 more sections
  • Runs Python scripts from its folder; calls python3; needs FIRECRAWL_API_KEY

What it does

Universal Scraping Architect is an agent skill from alirezarezvani/claude-skills. Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/firecrawl-technical-guide.md`, `references/local-extraction-patterns.md` and `references/scraping-ethics-security.md`).

It sits in Data & Analytics, covering Web scraping. It works with Python and Firecrawl. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • Document extraction
  • Building validation-heavy data pipelines using Firecrawl
  • Local Python scripts

Example prompts

  • “/universal-scraping-architect”

Requirements

  • Python 3
  • A credential in FIRECRAWL_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Route the Approach: Explicitly state whether Firecrawl or Local Python is being used and why.
  2. Track Budgets: Estimate Firecrawl API quotas or LLM token context limits before executing large jobs.
  3. Extract Safely: Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable…
  4. Validate & Clean: Run python3 scripts/validate_extraction.py extracted_output.json --json on every extraction result before delivering it…
  5. Format: Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text.

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FIRECRAWL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Universal Scraping Architect loads about 1.1k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 45 tokens; SKILL.md has 558 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 558 words, ~1,147 tokens.

Download SKILL.mdSave it as .claude/skills/universal-scraping-architect/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
universal-scraping-architect
description
Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

Universal Scraping Architect

Design complete, robust data-extraction pipelines with intelligent routing, validation, and token-budget tracking — not brittle one-off scripts.

Dependency Notice: BYOK (Bring Your Own Key) pattern for Firecrawl; API keys must only be loaded via environment variables. Per-script dependencies:

ScriptDependenciesExact CLI
scripts/validate_extraction.pystdlib onlypython3 scripts/validate_extraction.py output.json --json
scripts/firecrawl_example.pyfirecrawl, requests (template; --sample runs offline)python3 scripts/firecrawl_example.py --sample
scripts/local_bs4_example.pybeautifulsoup4, pandas (template; --sample runs offline)python3 scripts/local_bs4_example.py --sample

Before Starting

Check for context first: If project-context.md exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code.

How This Skill Works

This skill supports 3 extraction modes based on intelligent routing:

Mode 1: API-Driven (Firecrawl)

Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain.

Mode 2: Local Python (Traditional)

Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill.

Mode 3: Hybrid Pipeline

Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving.

The Extraction Pipeline

When executing a scraping task, always follow this sequence:

  1. Route the Approach: Explicitly state whether Firecrawl or Local Python is being used and why.
  2. Track Budgets: Estimate Firecrawl API quotas or LLM token context limits before executing large jobs.
  3. Extract Safely: Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. Start from the editable runner templates — scripts/firecrawl_example.py (Mode 1) or scripts/local_bs4_example.py (Mode 2); run each with --sample first to see the expected summary shape without network access.
  4. Validate & Clean: Run python3 scripts/validate_extraction.py extracted_output.json --json on every extraction result before delivering it. It exits 0 only on {"status": "ok"}; warning (empty output) or error (malformed JSON) exit 1 — fix and re-extract, never ship unvalidated data. Beyond this structural gate, also check required fields and duplicates against the pipeline spec before delivering.
  5. Format: Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text.
Show full SKILL.md (206 more words)Show less

Proactive Triggers

Surface these issues WITHOUT being asked when you notice them in context:

  • Hardcoded API Keys → Flag immediately and rewrite to use os.getenv('FIRECRAWL_API_KEY').
  • Private Data Leakage → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python).
  • Missing Pagination → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing.

Output Artifacts

When you ask for...You get...
"Scrape this site"A fully validated Python extraction script with routing logic and error handling.
"Get data from this table"A clean CSV/JSON dataset with a summary log of row counts and empty values.
"Crawl these docs"A Markdown deliverable chunked for LLM token limits.

Anti-Patterns

  • Brittle Selectors: Never use highly nested CSS selectors (e.g., div > span > ul > li:nth-child(3)). Use data attributes or robust structural anchors.
  • Ignoring Etiquette: Never scrape without checking robots.txt or implementing sensible rate limits.
  • No Validation: Never blindly write scraped data to a file without checking if the array is empty or missing critical keys.
  • data-cleaning: Use when the scraped data requires complex statistical normalization or deduplication.
  • browser-automation: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient.

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in engineering/universal-scraping-architect/skills/universal-scraping-architect of alirezarezvani/claude-skills.

  • SKILL.md
  • references/firecrawl-technical-guide.md
  • references/local-extraction-patterns.md
  • references/scraping-ethics-security.md
  • requirements.txt
  • scripts/firecrawl_example.py
  • scripts/local_bs4_example.py
  • scripts/validate_extraction.py

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Universal Scraping Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Universal Scraping Architect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Universal Scraping Architect this skillalirezarezvani/claude-skills28k—~1.1kAutomated safety check: PassMIT
Firecrawl Install Authjeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Firecrawl SDK Patternsjeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Firecrawl Page Scrape Integrationfirecrawl/firecrawl190k1 repos~944Automated safety check: PassISC
Firecrawl Interact Integrationfirecrawl/firecrawl190k1 repos~731Automated safety check: PassISC
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Firecrawl Install Auth

    jeremylongshore/tons-of-skills-marketplace

    Install the current Firecrawl Node or Python SDK, configure Cloud authentication, and verify package provenance and secret injection.

    2.8k GitHub stars~1.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Firecrawl SDK Patterns

    jeremylongshore/tons-of-skills-marketplace

    Build a typed, testable Firecrawl v2 adapter for Node or Python with current methods, explicit options, errors, pagination, and dependency control.

    2.8k GitHub stars~1.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.

    190k GitHub starsUsed in 1 repo~944 tokens
    Data & AnalyticsAuto-check passed
  • Guides adding Firecrawl's /interact endpoint to product code for pages that need clicks, forms, pagination or logged-in flows beyond plain scraping.

    190k GitHub starsUsed in 1 repo~731 tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 11 days ago
    Data & AnalyticsAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Works with

Questions about Universal Scraping Architect

What does Universal Scraping Architect do?

A skill your agent uses for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts. Universal Scraping Architect is an agent skill from alirezarezvani/claude-skills. Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

When should I use Universal Scraping Architect?

Universal Scraping Architect fits situations like: document extraction; building validation-heavy data pipelines using Firecrawl; local Python scripts.

How do I install Universal Scraping Architect in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill universal-scraping-architect -a claude-code`. Or copy the skill folder (engineering/universal-scraping-architect/skills/universal-scraping-architect in alirezarezvani/claude-skills) into .claude/skills/universal-scraping-architect in your project. Claude Code loads it when a task matches its description.

How do I install Universal Scraping Architect in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill universal-scraping-architect -a codex`. Or copy the skill folder (engineering/universal-scraping-architect/skills/universal-scraping-architect in alirezarezvani/claude-skills) into .agents/skills/universal-scraping-architect in your project. Codex loads it when a task matches its description.

Can I use Universal Scraping Architect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill universal-scraping-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/universal-scraping-architect, .gemini/skills/universal-scraping-architect, .github/skills/universal-scraping-architect and .opencode/skills/universal-scraping-architect in your project.

What does Universal Scraping Architect need to run?

Going by SKILL.md and its folder, Universal Scraping Architect needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named FIRECRAWL_API_KEY. Our summary lists: Python 3; A credential in FIRECRAWL_API_KEY.

Does Universal Scraping Architect access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Universal Scraping Architect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Universal Scraping Architect use?

Universal Scraping Architect is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Universal Scraping Architect use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Universal Scraping Architect?

Skills that share tags, products or a category with Universal Scraping Architect: Firecrawl Install Auth (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Firecrawl SDK Patterns (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Firecrawl Page Scrape Integration (firecrawl/firecrawl, 190k stars) and Firecrawl Interact Integration (firecrawl/firecrawl, 190k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Universal Scraping Architect?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,938 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.