Topic · Data & Analytics
Best web scraping skills, page 9
Web scraping skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 385 | A skill your agent uses when the user wants to interact with LinkedIn — extract post data (content, author, reactions, comments), scrape a member profile (name, headline, experience, education… | browsing-skills/ | 117 | — | ~575 | Automated safety check: Pass | MIT | 4 mo ago |
| 386 | A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows. | alirezarezvani/ | 28k | — | ~3.4k | Automated safety check: Notes | MIT | 1 mo ago |
| 387 | A skill your agent uses when a user needs a screenshot, website thumbnail, full-page capture, or PDF of a public HTTP(S) webpage saved as a local artifact through Latchshot, including report, QA… | github/ | 40k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 388 | Build production-ready Bright Data integrations with best practices baked in. | brightdata/ | 264 | 1 repo | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 389 | 389.Yc Jobs Scraper Scrape daily job listings from YCombinator's Workatastartup platform without duplicates. | Varnan-Tech/ | 674 | — | ~772 | Automated safety check: Pass | MIT | 1 mo ago |
| 390 | This skill should be used when the user asks to "find flights", "compare flights", "analyze flight search results", "plan a flight itinerary", "plan a complete trip", "plan an event trip", or create… | apify/ | 265 | — | ~1.1k | Automated safety check: Pass | MIT | 16 days ago |
| 391 | 391.RAG Pipeline Build a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the… | brightdata/ | 264 | — | ~1.9k | Automated safety check: Pass | MIT | yesterday |
| 392 | 392.Web Scraping Web scraping with anti-bot bypass, content extraction, undocumented APIs and poison pill detection. | nicepkg/ | 194 | 1 repo | ~5k | Automated safety check: Pass | No licence | 7 mo ago |
| 393 | 393.Browser Interactive Chromium sessions for reading and driving real web pages: open a URL, read the page as text, click and type, and read the page's own console and network history. | iii-hq/ | 113 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | today |
| 394 | 394.Venice Augment Venice augmentation endpoints for agent pipelines. An agent skill from veniceai/skills. | veniceai/ | 143 | — | ~2.5k | Automated safety check: Pass | MIT | 3 days ago |
| 395 | Fetch live Etsy search listing rows for a keyword, market phrase, or category via Apify Actor publicrecords/etsy-search-scraper (MCP). | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 396 | Read Etsy shop sales counters, deltas, and breakout flags from Apify Actor publicrecords/etsy-shop-velocity (MCP panel snapshot). | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 397 | Use Xquik for X data workflows: tweet search, user lookup, follower export, media downloads, monitors, webhooks, REST API, MCP, SDK setup, and approval-gated account actions. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 398 | 398.Geo Audit Runs a one-command technical GEO audit on any domain. An agent skill from vellum-ai/vellum-assistant. | vellum-ai/ | 1.4k | — | ~2k | Automated safety check: Pass | MIT | today |
| 399 | Expert blueprint for roguelikes including procedural generation (Walker method, BSP rooms), permadeath with meta-progression (unlock persistence), run state vs meta state separation, seeded RNG… | thedivergentai/ | 816 | — | ~4.4k | Automated safety check: Pass | LGPL-3.0 | 29 days ago |
| 400 | Drive a live browser on a scraped page: click, fill forms, log in, paginate, infinite-scroll. | firecrawl/ | 117 | — | ~738 | Automated safety check: Pass | ISC | today |
| 401 | A skill your agent uses when a user provides ChatGPT web AI-search keywords, repeat count, target entity, entity type, OpenCLI profile, and crawl interval preference, then needs repeated crawls… | yaojingang/ | 871 | — | ~456 | Automated safety check: Pass | MIT | 8 days ago |
| 402 | 402.Douyin Provide authenticated cookies for Douyin (抖音/TikTok China) — two providers: douyin (www.douyin.com for scraping) and douyin-live (live.douyin.com for livestream). | sigcli/ | 293 | — | ~696 | Automated safety check: Notes | MIT | 11 days ago |
| 403 | 403.Tiktok Provide authenticated cookies for TikTok (www.tiktok.com). An agent skill from sigcli/sigcli. | sigcli/ | 293 | — | ~479 | Automated safety check: Pass | MIT | 11 days ago |
| 404 | Complete knowledge domain for Firecrawl v2 API - web scraping and crawling that converts websites into LLM-ready markdown or structured data. | ynulihao/ | 617 | — | ~4.1k | Automated safety check: Notes | MIT | 7 mo ago |
| 405 | 405.Kol Engager Icp Find ICP-fit leads from KOL audiences on LinkedIn. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~1.7k | Automated safety check: Notes | MIT | today |
| 406 | Scrapes LinkedIn job postings using the JobSpy library (python-jobspy). | gooseworks-ai/ | 1.2k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | today |
| 407 | Find warm leads by searching LinkedIn for pain-language posts — the frustrations, complaints, and operational struggles your ICP talks about publicly. | gooseworks-ai/ | 1.2k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | today |
| 408 | Extract commenters from LinkedIn posts via Apify. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~722 | Automated safety check: Pass | MIT | today |
| 409 | Search LinkedIn posts by keywords, sorted by engagement or date. | gooseworks-ai/ | 1.2k | 1 repo | ~1.3k | Automated safety check: Notes | MIT | today |
| 410 | Scrape recent posts from LinkedIn profiles using Apify. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~495 | Automated safety check: Pass | MIT | today |
| 411 | 411.Meta Ad Scraper Scrape competitor ads from Meta's Ad Library (Facebook, Instagram, Messenger, Threads, WhatsApp). | gooseworks-ai/ | 1.2k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 412 | 412.Review Scraper Scrape product reviews from G2, Capterra, and Trustpilot using Apify. | gooseworks-ai/ | 1.2k | 1 repo | ~573 | Automated safety check: Pass | MIT | today |
| 413 | Scrape product reviews from G2, Capterra, and Trustpilot using Apify. | gooseworks-ai/ | 1.2k | 1 repo | ~882 | Automated safety check: Pass | MIT | today |
| 414 | Find TikTok influencers using Apify's Influencer Discovery Agent. | gooseworks-ai/ | 1.2k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | today |
| 415 | Search and scrape Twitter/X posts using Apify. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~805 | Automated safety check: Pass | MIT | today |
| 416 | 416.Cnki Crawler Crawl CNKI journal paper metadata with the local Python crawler and save to PostgreSQL. | Drchronx/ | 137 | — | ~1.4k | Automated safety check: Notes | Unknown | 4 mo ago |
| 417 | A skill your agent uses for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts. | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 418 | 418.Browser Use A skill your agent uses to drive Wisp Browser Runtime sessions (shared daily Chrome or workspace Chrome) — open pages, read them, click, fill and submit forms, navigate, switch tabs, or scrape… | xuzhougeng/ | 1k | — | ~2.7k | Automated safety check: Pass | AGPL-3.0 | today |
| 419 | 419.Deep Research Multi-source deep research using firecrawl and exa MCPs. An agent skill from aAAaqwq/AGI-Super-Team. | aAAaqwq/ | 105 | 5 repos | ~1.1k | Automated safety check: Pass | MIT | today |
| 420 | Download genome assemblies, gene records, and ortholog data from NCBI using the modern Datasets v2 CLI (replaces assemblysummary.txt scraping and many EFetch workflows). | GPTomics/ | 1.2k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 421 | Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store. | apify/ | 265 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 422 | 422.Scrape Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (bdata scrape). | brightdata/ | 264 | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 423 | 423.Search Search the web via the Bright Data CLI — bdata search for Google/Bing/Yandex SERP, bdata discover for intent-ranked semantic results. | brightdata/ | 264 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 424 | 424.SEO Audit and improve website discoverability across traditional search, answer engines, and generative search. | magnus919/ | 116 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 425 | Collect Bilibili video and creator evidence with MediaCrawler through search, exact BV detail, comments, dynamics, contacts, and optional media. | tsingyuai/ | 2k | — | ~356 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 426 | Collect Douyin competitive evidence with MediaCrawler using keyword search, exact video detail, comments, media, and creator profiles. | tsingyuai/ | 2k | — | ~381 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 427 | Collect Kuaishou competitive evidence with MediaCrawler using keyword search, exact video detail, comments, media, and creator profiles. | tsingyuai/ | 2k | — | ~338 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 428 | Collect Baidu Tieba thread and user evidence with MediaCrawler using keyword or bar discovery, exact thread detail, replies, and creator pages. | tsingyuai/ | 2k | — | ~338 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 429 | Collect Weibo posts and creator evidence with MediaCrawler using search, exact post IDs, comments, optional media, and creator IDs. | tsingyuai/ | 2k | — | ~341 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 430 | Collect Zhihu answers, articles, videos, comments, and creator evidence with MediaCrawler through search and exact URLs. | tsingyuai/ | 2k | — | ~350 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 431 | 任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub… | affaan-m/ | 276k | — | ~377 | Automated safety check: Pass | MIT | 4 days ago |
| 432 | A skill your agent uses when the user wants to integrate with the X (Twitter) API via Xquik to search tweets, look up user profiles, extract followers, run giveaway draws, monitor accounts, or… | davila7/ | 32k | — | ~2.3k | Automated safety check: Pass | MIT | today |