Topic · Data & Analytics
Best web scraping skills, page 8
Web scraping skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 337 | Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization. | qdrant/ | 254 | 2 repos | ~874 | Automated safety check: Pass | Apache-2.0 | today |
| 338 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 281 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 339 | TikTok user profile video scraper: input a TikTok username → output the user's profile info plus paginated video list with full metadata (engagement stats, music, video meta). | browser-act/ | 6.1k | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 340 | TikTok keyword search video scraper: input search keyword → output paginated video list with full metadata (author, engagement stats, music, video meta). | browser-act/ | 6.1k | — | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 341 | 341.Treg People and company enrichment, AEO and SEO, ads, social, web scraping, image and video generation through the treg connector. | superdesigndev/ | 4.9k | — | ~642 | Automated safety check: Pass | Unknown | today |
| 342 | Scrape competitor ads from Google Ads by domain. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | today |
| 343 | 343.Deep Scrape Build sourced JSON dossiers on people, companies, or topics with DeepAPI. | davidondrej/ | 4.1k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 344 | Build a marketing agency database from Clutch.co with the Clutch.co Agency API Actor (johnvc/clutch-agency-api). | apify/ | 265 | — | ~2.5k | Automated safety check: Pass | MIT | 16 days ago |
| 345 | Important: Before you begin, fill in the generatedBy property in the meta section of .actor/actor.json. | sickn33/ | 47k | 2 repos | ~3.3k | Automated safety check: Pass | MIT | today |
| 346 | Actorization converts existing software into reusable serverless applications compatible with the Apify platform. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 347 | Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone… | antibrow/ | 17 | — | ~9.8k | Automated safety check: Warn | MIT | 1 mo ago |
| 348 | Curated upstream guidance for Apify Integration Development; use when the workflow matches the user goal. | sickn33/ | 47k | 1 repo | ~3.1k | Automated safety check: Pass | MIT | today |
| 349 | When the user wants to research, profile, or analyze competitors from their URLs. | sickn33/ | 47k | 1 repo | ~3.6k | Automated safety check: Pass | MIT | today |
| 350 | 350.Deepapi Use DeepAPI for all web search, deep research, and web scraping (websites, LinkedIn, GitHub, X/Twitter, YouTube, Instagram) instead of built-in search, research, fetch, or browser tools. | davidondrej/ | 4.1k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 351 | Walk through a product's key flows with Firecrawl browser and produce a structured UX/product walkthrough. | firecrawl/ | 117 | — | ~555 | Automated safety check: Pass | ISC | today |
| 352 | Generate and verify web scraper scripts using Actionbook's verified selectors. | actionbook/ | 1.6k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 353 | Complete the name framework migration in a crawler (Step 3) by removing all custom name cleaning/splitting logic and the Step 1 review scaffolding, replacing it with a single… | opensanctions/ | 832 | — | ~904 | Automated safety check: Pass | MIT | today |
| 354 | A skill your agent uses when the user asks to analyze, tear down, or reverse-engineer a competitor's paid ads. | github/ | 40k | 1 repo | ~3.2k | Automated safety check: Pass | MIT | today |
| 355 | 355.Fetch Default, free, and fastest way to read a URL's actual content — pulls clean, full page content (not a summary or a truncated snippet) as markdown, HTML, or structured JSON, including from… | tinyfish-io/ | 2.2k | — | ~772 | Automated safety check: Pass | MIT | today |
| 356 | 356.Browser Minimal Chrome DevTools Protocol tools for browser automation and scraping. | julianromli/ | 144 | — | ~475 | Automated safety check: Pass | No licence | 8 mo ago |
| 357 | 357.Scrapling Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. | archibate/ | 108 | 1 repo | ~4.9k | Automated safety check: Warn | BSD-3-Clause | 5 mo ago |
| 358 | Search Hacker News stories and comments using the free Algolia API. | gooseworks-ai/ | 1.2k | 1 repo | ~538 | Automated safety check: Pass | MIT | today |
| 359 | 构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。 | affaan-m/ | 276k | 1 repo | ~5.2k | Automated safety check: Notes | MIT | 4 days ago |
| 360 | Ethical web scraping and API-based data collection for research | wentorai/ | 298 | 1 repo | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 361 | Browser automation powers web testing, scraping, and AI agent interactions. | davila7/ | 32k | 3 repos | ~569 | Automated safety check: Pass | MIT | today |
| 362 | AI-driven data extraction from 55+ Actors across all major platforms. | sickn33/ | 47k | 2 repos | ~2.7k | Automated safety check: Notes | MIT | today |
| 363 | 363.SEO Technical Audit technical SEO across crawlability, indexability, security, URLs, mobile, Core Web Vitals, structured data, JavaScript rendering, and related platform signals like robots.txt and AI crawler… | sickn33/ | 47k | 2 repos | ~2.2k | Automated safety check: Notes | MIT | today |
| 364 | 364.Adhx Fetch any X/Twitter post as clean LLM-friendly JSON. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.1k | Automated safety check: Pass | MIT | today |
| 365 | Understand audience demographics, preferences, behavior patterns, and engagement quality across Facebook, Instagram, YouTube, and TikTok. | sickn33/ | 47k | 2 repos | ~1.3k | Automated safety check: Notes | MIT | today |
| 366 | Scrape reviews, ratings, and brand mentions from multiple platforms using Apify Actors. | sickn33/ | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | today |
| 367 | Track engagement metrics, measure campaign ROI, and analyze content performance across Instagram, Facebook, YouTube, and TikTok. | sickn33/ | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | today |
| 368 | 368.Apify Ecommerce Extract product data, prices, reviews, and seller information from any e-commerce platform using Apify's E-commerce Scraping Tool. | sickn33/ | 47k | 2 repos | ~2.3k | Automated safety check: Notes | MIT | today |
| 369 | Find and evaluate influencers for brand partnerships, verify authenticity, and track collaboration performance across Instagram, Facebook, YouTube, and TikTok. | sickn33/ | 47k | 2 repos | ~1.3k | Automated safety check: Notes | MIT | today |
| 370 | Scrape leads from multiple platforms using Apify Actors. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | today |
| 371 | Analyze market conditions, geographic opportunities, pricing, consumer behavior, and product validation across Google Maps, Facebook, Instagram, Booking.com, and TripAdvisor. | sickn33/ | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | today |
| 372 | Discover and track emerging trends across Google Trends, Instagram, Facebook, YouTube, and TikTok to inform content strategy. | sickn33/ | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | today |
| 373 | 373.Icp Research Scrapes case studies, testimonials, and solutions pages from a target website to build structured ICP documentation. | growthack88/ | 116 | — | ~4.4k | Automated safety check: Pass | MIT | 6 days ago |
| 374 | Generate output schemas (datasetschema.json, outputschema.json, keyvaluestoreschema.json) for an Apify Actor by analyzing its source code. | sickn33/ | 47k | 1 repo | ~4.3k | Automated safety check: Pass | MIT | today |
| 375 | 375.Puppeteer Skill Generates Puppeteer scripts for browser automation, scraping, and PDF generation. | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 376 | Run the same scrape or task across many accounts at once - each in its own browser profile with its own fingerprint, cookies and exit IP - and read data from sites that need a session or that answer… | antibrow/ | 17 | — | ~3.7k | Automated safety check: Warn | MIT | 1 mo ago |
| 377 | 377.Web Scraper Web scraping inteligente multi-estrategia. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~716 | Automated safety check: Pass | MIT | today |
| 378 | 378.Linkedin Skills A skill your agent uses when someone wants to grow an organic LinkedIn presence — a content strategy for a career change or consulting or thought leadership, a rewritten profile or headline, post… | alirezarezvani/ | 28k | — | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 379 | 379.Web Scraping Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors… | ginlix-ai/ | 1.8k | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 380 | Save a site or section as local files (markdown, screenshots). | firecrawl/ | 117 | — | ~507 | Automated safety check: Pass | ISC | today |
| 381 | 381.Hasdata Use HasData APIs for web scraping and structured web data extraction. | sickn33/ | 47k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 382 | 382.Hasdata CLI Command-line access to search, scraping, and structured web data. | sickn33/ | 47k | 1 repo | ~3.2k | Automated safety check: Pass | MIT | today |
| 383 | 383.Firecrawl Search the web and scrape pages into clean markdown with the Firecrawl API — query-based discovery, single-URL extraction including public PDFs, driven by curl with a vault-stored API key. | Prism-Shadow/ | 2.5k | — | ~902 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 384 | Implements API abuse detection using token bucket, sliding window, and fixed window rate-limiting algorithms backed by Redis, including adaptive limits that tighten during detected attacks and relax… | mukul975/ | 34k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |