Topic · Data & Analytics
Best web scraping skills, page 10
Web scraping skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 433 | 433.Ops Leadgen OPS on-demand: This skill should be used when the user asks to "leadgen drafts", "cold email approve"… | Lifecycle-Innovations-Limited/ | 540 | — | ~990 | Automated safety check: Notes | MIT | yesterday |
| 434 | Generate a personalized outbound message (LinkedIn DM or cold email) for a single prospect by combining a template with the lead's profile, company, and recent signals. | Othmane-Khadri/ | 317 | — | ~1.3k | Automated safety check: Notes | MIT | 1 mo ago |
| 435 | Pull the audience that engaged with a LinkedIn post — likers (reactors) and commenters — via Unipile, dedupe across endpoints, and persist them as a result set ready for qualification or campaign… | Othmane-Khadri/ | 317 | — | ~1.4k | Automated safety check: Notes | MIT | 1 mo ago |
| 436 | 436.Interview Prep Generate a structured interview preparation guide for any company by scraping real candidate experiences from Glassdoor, Blind, and Reddit in real time using parallel TinyFish agents. | tinyfish-io/ | 2.2k | — | ~2.3k | Automated safety check: Pass | MIT | 7 days ago |
| 437 | Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy… | grafana/ | 279 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 438 | 438.Yelp Search Search Yelp for local businesses, get contact info, ratings, and hours. | letta-ai/ | 147 | — | ~1.4k | Automated safety check: Notes | MIT | 7 days ago |
| 439 | 439.Wps Doc Scraper Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown. | daymade/ | 1.4k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 440 | Non-testing browser automation - web scraping, form filling, screenshot capture, PDF generation, workflow automation. | ynulihao/ | 617 | — | ~2.2k | Automated safety check: Pass | No licence | 7 mo ago |
| 441 | Debug Bright Data Scraping Browser sessions using the Browser Sessions API. | brightdata/ | 264 | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 442 | 442.Brightdata SDK Web data extraction and discovery using the Bright Data Python SDK. | brightdata/ | 264 | — | ~5.2k | Automated safety check: Pass | MIT | yesterday |
| 443 | 443.Scraper Builder Build production-ready web scrapers for any website using Bright Data infrastructure. | brightdata/ | 264 | — | ~7.2k | Automated safety check: Pass | MIT | yesterday |
| 444 | Monitor competitor pricing pages via live web scrape and Web Archive snapshots. | gooseworks-ai/ | 1.2k | 1 repo | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 445 | Extract leads from competitor product activity — Product Hunt commenters/upvoters, HN posts about competitors, case studies, testimonials, tech press, and switching signals. | gooseworks-ai/ | 1.2k | 1 repo | ~3k | Automated safety check: Notes | MIT | yesterday |
| 446 | Monitor web sources for Series A-C funding announcements. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 447 | Prepare for investor calls by pulling upcoming meetings from Google Calendar, deeply researching each investor and their firm (website scraping, portfolio analysis, thesis extraction), checking for… | gooseworks-ai/ | 1.2k | 1 repo | ~3.8k | Automated safety check: Pass | MIT | yesterday |
| 448 | 448.Job Scraper Search for job postings across LinkedIn and Indeed. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~2.6k | Automated safety check: Notes | MIT | yesterday |
| 449 | Track what key opinion leaders (KOLs) in your space are posting on LinkedIn and Twitter/X. | gooseworks-ai/ | 1.2k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 450 | Lead qualification engine with conversational intake. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~3.8k | Automated safety check: Pass | MIT | yesterday |
| 451 | Generates Instagram-ready product reels from any e-commerce product page URL. | gooseworks-ai/ | 1.2k | 1 repo | ~2.3k | Automated safety check: Notes | MIT | yesterday |
| 452 | Scrape G2, Capterra, and Trustpilot reviews for your product and competitors, then extract recurring themes, objections, proof points, and exact customer language for use in messaging. | gooseworks-ai/ | 1.2k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 453 | Pull real SEO metrics for any domain using Apify scrapers for Semrush and Ahrefs data. | gooseworks-ai/ | 1.2k | 1 repo | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 454 | 454.Signal Scanner Detect buying signals across TAM companies and watchlist personas. | gooseworks-ai/ | 1.2k | 1 repo | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 455 | 455.Web Scraping Scrape websites, extract structured data, and automate browsers. | gooseworks-ai/ | 1.2k | 1 repo | ~4.6k | Automated safety check: Pass | MIT | yesterday |
| 456 | 456.Firecrawl Crawl Bulk extract content from an entire website or site section. | aiskillstore/ | 430 | 3 repos | ~673 | Automated safety check: Pass | No licence | yesterday |
| 457 | 457.Firecrawl Web search and scraping via Firecrawl API. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | 1 repo | ~250 | Automated safety check: Notes | No licence | 7 mo ago |
| 458 | 458.Firecrawl Map Discover and list a site's URLs, with search filtering. An agent skill from firecrawl/skills. | firecrawl/ | 116 | 1 repo | ~412 | Automated safety check: Pass | ISC | 2 days ago |
| 459 | Use before adding or changing Python code that calls a third-party HTTP API from PostHog (a vendor REST call, a vendor SDK client, a scraping or enrichment service), and before adding or changing a… | PostHog/ | 40k | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 460 | Scrape Reddit with the harshmaur/reddit-scraper Apify Actor: keyword search across all of Reddit or inside one subreddit, full subreddit listings, post permalinks with complete comment threads, user… | apify/ | 264 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 461 | 461.Kb Refresh Add new sources to your knowledge base or re-scrape existing ones to pick up changes. | techwolf-ai/ | 132 | 1 repo | ~1.3k | Automated safety check: Pass | MIT | 9 days ago |
| 462 | 462.Scrape Scrape any webpage as clean markdown via Bright Data Web Unlocker API. | davila7/ | 32k | — | ~392 | Automated safety check: Pass | MIT | yesterday |
| 463 | Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. | davepoon/ | 3.6k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 464 | 464.Browser Scrape DEPRECATED in v0.2.0 -- use browser-extract instead; this is a thin shim for backward compatibility, removed in v0.3.0 | ruvnet/ | 74k | — | ~405 | Automated safety check: Notes | MIT | yesterday |
| 465 | 小红书作品爬取工具。根据关键词爬取小红书热门作品数据,支持按日期范围、排序方式筛选,结果以结构化表格展示。当用户需要爬取小红书作品、查询小红书热门内容、搜索小红书爆款笔记时使用。触发词:小红书爬取、小红书作品、小红书爆款、小红书搜索、小红书热门、小红书笔记查询。 | redfox-data/ | 425 | — | ~1.5k | Automated safety check: Pass | No licence | yesterday |
| 466 | Build GitHub Copilot workflows with Xquik X API SDKs, REST endpoints, hosted Apify Actor runs, MCP tools, TweetClaw OpenClaw plugin installs, signed webhooks, tweet search, user lookup, follower… | github/ | 40k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 467 | 467.Barchart Quotes, fundamentals, insider, analyst ratings, earnings estimates, financial summary, income/balance/cashflow detail y company profile (delayed 15-20min). | gauss314/ | 246 | — | ~5.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 468 | Run Xquik's Apify Actor for X followers, following, verified audiences, lists, communities, and overlap research. | Varnan-Tech/ | 674 | — | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 469 | Run Xquik's Apify Actor for X searches, posts, timelines, conversations, lists, articles, and engagement research. | Varnan-Tech/ | 674 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 470 | A skill your agent uses when writing Playwright automation code, building web scrapers, or creating E2E tests - provides best practices for selector strategies, waiting patterns, and robust… | ed3dai/ | 250 | — | ~3.6k | Automated safety check: Pass | No licence | 1 mo ago |
| 471 | Run ad-hoc web research on a prospect (company URL or person) — Firecrawl scrapes the marketing site, key pages (pricing, careers, blog), and surfaces a brief. | Othmane-Khadri/ | 317 | — | ~490 | Automated safety check: Notes | MIT | 1 mo ago |
| 472 | Pull competitive intelligence on a competitor — pricing, positioning, recent changes, customer stories — via the competitive-intel CLI. | Othmane-Khadri/ | 317 | — | ~445 | Automated safety check: Notes | MIT | 1 mo ago |
| 473 | 473.Deep Research Produce cited research reports from multiple web sources using firecrawl and exa MCP tools — plan sub-questions, search and deep-read sources, then synthesize findings with inline citations and… | affaan-m/ | 275k | — | ~1.5k | Automated safety check: Warn | MIT | 4 days ago |
| 474 | 474.Archive Crawler Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump). | inbrainfun/ | 142 | 1 repo | ~2.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 475 | Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser. | firecrawl/ | 116 | — | ~565 | Automated safety check: Pass | ISC | 2 days ago |
| 476 | Analyze competitor strategies, content, pricing, ads, and market positioning across Google Maps, Booking.com, Facebook, Instagram, YouTube, and TikTok. | majiayu000/ | 666 | 4 repos | ~1.3k | Automated safety check: Notes | MIT | yesterday |
| 477 | 477.Cf Crawl Crawl entire websites using Cloudflare Browser Rendering /crawl API. | davila7/ | 32k | — | ~2.6k | Automated safety check: Notes | MIT | yesterday |
| 478 | X API & Twitter scraper skill for AI coding agents. An agent skill from davila7/claude-code-templates. | davila7/ | 32k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 479 | A skill your agent uses to report Omni-Channel Pending Service Routing usage without changing the org: query the current PendingServiceRouting count, compare it with an admin-supplied maximum or a… | forcedotcom/ | 1.1k | — | ~964 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 480 | Extracts Feishu/Lark Docs, Wiki, Sheets, and Minutes (妙记) transcripts into faithful local Markdown via the lark-cli API — no LLM paraphrasing, browser-DOM fallback when lark-cli can't reach content. | daymade/ | 1.4k | — | ~6.9k | Automated safety check: Pass | MIT | yesterday |