Topic · Data & Analytics
Best web scraping skills, page 4
Web scraping skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Scraper Add, debug, or modify price scrapers for e-commerce sites (Shopee, Lazada, Tiki, etc.). | duyet/ | 139 | — | ~451 | Automated safety check: Pass | MIT | 2 days ago |
| 146 | 146.Gog Pamphlet Builds the printable pamphlet panels that go inside a GOG backup disc case — scrapes the store page, generates the QR badge, writes the per-title JSON and runs buildset.py to emit 483x629 SVG and… | east35/ | 121 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | 11 days ago |
| 147 | Refactor the title, description and coverage frequency of a legislature/parliament PEP dataset .yml into the house style. | opensanctions/ | 832 | — | ~953 | Automated safety check: Pass | MIT | today |
| 148 | Exports the article list and original article text from a Tencent ima knowledge base through a logged-in Chrome session, using browser automation. | zj-unicom-ai/ | 358 | — | ~1k | Automated safety check: Pass | MIT | today |
| 149 | 149.Web Scrape Intelligent web scraper with content extraction, multiple output formats, and error handling | 21pounder/ | 120 | 1 repo | ~1.5k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 150 | 150.Apify Collect Collect fresh job-posting URLs into ~/.dear-hiring-manager/urls.txt by running an Apify job scraper — the discovery source for /batch. | extrasmall0/ | 111 | — | ~1.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 151 | Turn a Zillow listing into a cinematic room-by-room walkthrough video (Apify scrape → Higgsfield image-to-video → ffmpeg stitch) to sell to real estate agents | charlesdove977/ | 163 | — | ~1.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 152 | Use this in the AI blogger crawler project when Douyin daily creator updates should use the Codex Chrome plugin instead of the local web-access CDP proxy to collect creator-page cards, selected… | dragon-hh/ | 106 | — | ~765 | Automated safety check: Pass | No licence | 1 mo ago |
| 153 | 153.Web Unblocker Bypasses anti-bot protections using Oxylabs Web Unblocker, an AI-powered proxy that handles fingerprinting, JavaScript rendering, and retries automatically. | oxylabs/ | 875 | — | ~865 | Automated safety check: Pass | MIT | 7 days ago |
| 154 | 154.Scrapingbee CLI Fetch and read any web page, search the web, crawl a site, or pull structured data out of pages. | ScrapingBee/ | 108 | — | ~3.2k | Automated safety check: Notes | MIT | today |
| 155 | Find leads by scraping engagers from a competitor's top LinkedIn posts. | gooseworks-ai/ | 1.2k | 1 repo | ~1.8k | Automated safety check: Notes | MIT | today |
| 156 | 156.Xcrawl Scrape A skill your agent uses for XCrawl scrape tasks, including single-URL fetch, format selection, sync or async execution, and JSON extraction with prompt or jsonschema. | xcrawl-api/ | 449 | — | ~2.2k | Automated safety check: Pass | No licence | 6 mo ago |
| 157 | 157.Firecrawl Scrape Read a known webpage or execute a discovered workflow or data-provider capability. | firecrawl/ | 116 | — | ~1.8k | Automated safety check: Pass | ISC | yesterday |
| 158 | 158.Autobahn Carve guardrail-adjacent items out of scope with safe alternatives before risk-adjacent work starts, then run the safe remainder at full strength in a fresh subagent that only ever sees the carved… | LilMGenius/ | 1.1k | — | ~2.1k | Automated safety check: Pass | MIT | 6 days ago |
| 159 | 159.Fix Scraper Diagnose and fix a broken scraper mapping. An agent skill from RealEstateWebTools/property_web_scraper. | RealEstateWebTools/ | 132 | — | ~1k | Automated safety check: Pass | MIT | 2 mo ago |
| 160 | 160.Leetcode Py Generates Python LeetCode practice environments and manages a 307-problem catalog with the lcpy CLI. | wislertt/ | 142 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 161 | Guides designing and building an official Apify integration for a company's product: workflow-automation apps, AI agent plugins, AI framework packages or direct API clients. | apify/ | 2.4k | — | ~3.1k | Automated safety check: Pass | No licence | yesterday |
| 162 | Set up a recurring buying-signal detection pipeline that finds companies showing buying intent across three signal types — job postings (hiring for the persona), fundraising events (recent raises)… | apify/ | 264 | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | 15 days ago |
| 163 | 163.Supercompress Always-on context compression for Grok Build. An agent skill from Supercompress/Supercompress. | Supercompress/ | 106 | — | ~519 | Automated safety check: Pass | MIT | yesterday |
| 164 | Extracts Airbnb search results - listing id, name, coordinates, rating, price, photos and badges - from the page's own embedded search data, with pagination. | browser-act/ | 6.1k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 165 | 165.Chatgpt Search Search ChatGPT and extract the full response + hydration JSON that powers the UI. | SeifBenayed/ | 114 | — | ~1.7k | Automated safety check: Notes | MIT | 6 mo ago |
| 166 | Practical guidance for multi-source web research with Exa, the Firecrawl CLI and Reddit MCP tools: when to search, when to fetch and which tool to prefer. | malob/ | 463 | 1 repo | ~2.4k | Automated safety check: Pass | MIT | 4 days ago |
| 167 | Migrate ad-hoc name cleaning in a crawler to h.reviewnames (Step 1 of the name framework migration). | opensanctions/ | 832 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 168 | Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites. | firecrawl/ | 1.2k | — | ~760 | Automated safety check: Pass | MIT | 5 mo ago |
| 169 | Umbrella skill for a library of website-specific browsing skills. Use when the user's request targets one of these specific websites: <!… | browsing-skills/ | 116 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 170 | MUST USE when investigating performance issues on a ClickHouse-managed Postgres instance. | ClickHouse/ | 544 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 171 | 171.Kimi Webbridge Kimi Browser Extension(Kimi 浏览器扩展,原 Kimi WebBridge)lets AI control the user's real browser — navigate, click, type, read, screenshot, and interact with any website using the user's actual login… | mxyhi/ | 493 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 172 | 172.Web Scraper Scrape, crawl, and extract data from websites. An agent skill from shobcoder/shob. | shobcoder/ | 577 | — | ~700 | Automated safety check: Pass | MIT | 20 days ago |
| 173 | Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone… | antibrow/ | 932 | — | ~9.8k | Automated safety check: Warn | MIT | 1 mo ago |
| 174 | 174.Find Skills Run this BEFORE any package install (pip / npm / apt / brew / cargo / gem / go install) you would otherwise execute via the exec tool — including when the user asks for a deliverable that needs… | fastclaw-ai/ | 1.4k | — | ~2.3k | Automated safety check: Warn | Unknown | yesterday |
| 175 | Turns your latest successful /scrape run into a permanent browser skill with a script, a test and a fixture, so repeat scrapes run in about 200 ms. | garrytan/ | 136k | 1 repo | ~11k | Automated safety check: Notes | MIT | today |
| 176 | Autonomous cold email campaign launcher. An agent skill from growthenginenowoslawski/coldoutboundskills. | growthenginenowoslawski/ | 742 | — | ~2.9k | Automated safety check: Pass | MIT | 2 days ago |
| 177 | 177.Web Scraper 网页数据抓取专家。能提取网页正文、表格、列表、特定元素。适用于:竞品监控、价格追踪、结构化数据采集. An agent skill from charlie-chann/ai-agent-langgraph-main. | charlie-chann/ | 114 | — | ~431 | Automated safety check: Pass | No licence | 24 days ago |
| 178 | 178.Xcrawl Search A skill your agent uses for XCrawl search tasks, including keyword search request design, location and language controls, result analysis, and follow-up crawl or scrape planning. | xcrawl-api/ | 449 | — | ~1.1k | Automated safety check: Pass | No licence | 6 mo ago |
| 179 | Runs web-page JavaScript without a browser, in a V8 sandbox with a browser-environment shim, for scripts that only probe the environment and compute a result. | taxueseek/ | 185 | — | ~400 | Automated safety check: Pass | MIT | yesterday |
| 180 | Score and enrich a CSV of B2B leads using Apify Actors. An agent skill from apify/awesome-skills. | apify/ | 264 | — | ~4.4k | Automated safety check: Notes | Apache-2.0 | 15 days ago |
| 181 | 181.Web Access Routes every networked task to the lightest channel that reaches it: a direct search, a page fetch, or a persistent browser for logins and dynamic pages. | ZhanlinCui/ | 191 | — | ~1.4k | Automated safety check: Pass | MIT | 7 mo ago |
| 182 | Build high-quality literature reviews from a research topic using a 10-phase workflow. | Drchronx/ | 135 | — | ~2.5k | Automated safety check: Pass | Unknown | 4 mo ago |
| 183 | Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended)… | indranilbanerjee/ | 855 | 1 repo | ~3.9k | Automated safety check: Pass | MIT | 4 days ago |
| 184 | Finds a company's official website and social profiles from its name, or collects social links from a website URL, using BrowserAct templates run by a Python script. | browser-act/ | 6.1k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 185 | 185.Refactor Crawler Rewrite messy or AI-generated crawler code into clean, production-ready style that follows the zavod best practices. | opensanctions/ | 832 | — | ~1.1k | Automated safety check: Notes | MIT | today |
| 186 | 186.Tmux Remote control tmux sessions for interactive CLIs (python, gdb, etc.) by sending keystrokes and scraping pane output. | mitsuhiko/ | 3.2k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 187 | 187.Skill Gen Deprecated. An agent skill from crafter-station/skills. | crafter-station/ | 112 | — | ~5.7k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 188 | A skill your agent uses when a live page needs Chrome DevTools/CDP evidence: network failures, console errors, performance, DOM/CSS actionability, screenshots/PDF, cookies/storage… | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | 4 days ago |
| 189 | Run the same scrape or task across many accounts at once - each in its own browser profile with its own fingerprint, cookies and exit IP - and read data from sites that need a session or that answer… | antibrow/ | 932 | — | ~3.7k | Automated safety check: Warn | MIT | 1 mo ago |
| 190 | A skill your agent uses when improving internal link structure, anchor text, orphan pages, crawl depth, site architecture, or link equity flow. | ViryaZheng/ | 460 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 191 | Crawls, maps, and scrapes an entire site through the Firecrawl MCP server for broken-link checks and content inventories. | AgriciDaniel/ | 18k | 1 repo | ~2k | Automated safety check: Pass | MIT | 3 days ago |
| 192 | 192.Firecrawl Agent Autonomously navigate websites and extract structured data across pages. | firecrawl/ | 116 | — | ~1.2k | Automated safety check: Pass | ISC | yesterday |