Topic · Data & Analytics
Best web scraping skills, page 5
Web scraping skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store. | apify/ | 264 | 1 repo | ~2k | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 194 | Guides a policy notification agent through crawling government policy sources, saving them locally, matching them to enterprise profiles and generating notifications. | Serein-81/ | 148 | — | ~1.4k | Automated safety check: Notes | No licence | 4 mo ago |
| 195 | Triggered when the user asks for page-level interaction with a specific webpage: login, form filling, button clicks, pagination scraping, screenshots, downloading page images, handling CAPTCHAs or… | OpenLoaf/ | 108 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0 | 4 mo ago |
| 196 | A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies. | revfactory/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 197 | Find the TikTok photo-mode carousels (slideshows) that actually go viral in a niche and the accounts behind them, rank them by organic quality (save-rate, like-rate, boost detection) instead of raw… | nestyme/ | 151 | — | ~2.6k | Automated safety check: Notes | No licence | 8 days ago |
| 198 | 198.SEO Site Audit Run a complete technical SEO + AI-search audit of a website with crawlie. | spronta/ | 112 | — | ~966 | Automated safety check: Pass | Unknown | 2 mo ago |
| 199 | 199.Data Extractor Extract structured data from websites into CSV or JSON. An agent skill from hanzili/hanzi-browse. | hanzili/ | 177 | — | ~2.2k | Automated safety check: Pass | Unknown | 5 mo ago |
| 200 | 200.Twitter Crawler Twitter 推文爬取器 - 指定用户名爬取推文,保存为 Markdown 格式,支持自定义数量和字段. An agent skill from huangserva/servasyy_skills. | huangserva/ | 167 | — | ~672 | Automated safety check: Notes | No licence | 6 days ago |
| 201 | Searches Douyin (douyin.com) for videos by keyword and returns structured video data including author info, stats, cover, description, hashtags, and download URL. | browser-act/ | 6.1k | 1 repo | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 202 | Design and validate safe source connector manifests for public graduate-admission, policy, advisor, lab, and community sources. | Bubble252/ | 123 | — | ~379 | Automated safety check: Pass | MIT | 1 mo ago |
| 203 | 203.Tmux Remote control tmux sessions for interactive CLIs (python, gdb, etc.) by sending keystrokes and scraping pane output. | archibate/ | 108 | 1 repo | ~1.4k | Automated safety check: Pass | No licence | 5 mo ago |
| 204 | 204.News Fetch, search, and summarize financial news articles via the local scraper service (~/dev/scraper) at http://127.0.0.1:8089. | 9600dev/ | 131 | — | ~6.1k | Automated safety check: Pass | Unknown | 23 days ago |
| 205 | A skill your agent uses when the user says "enrich people who engaged with this post", "qualify post engagers with FullEnrich", "scrape and enrich LinkedIn post {URL}", "engagers from this post into… | Othmane-Khadri/ | 317 | — | ~1.5k | Automated safety check: Warn | MIT | 1 mo ago |
| 206 | Extract structured company lists from directories with Firecrawl. | firecrawl/ | 116 | — | ~557 | Automated safety check: Pass | ISC | yesterday |
| 207 | 207.Scrape Nowcoder 基于 CDP 原生 WebSocket 抓取牛客网面经文章。当用户说"抓牛客"、"爬牛客面经"、"nowcoder 抓取"、"抓取面经列表"时触发。通过 Chrome 调试端口直接连接已登录的浏览器会话,支持首页、话题、搜索分页和详情全文抓取,输出 Markdown。 | ranxi2001/ | 690 | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 208 | 208.Xrk Crawl 当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 webfetch 与 browser 工作流如何选型时使用。 | xrkseek/ | 140 | — | ~1.3k | Automated safety check: Pass | MIT | 7 days ago |
| 209 | Track which of your LinkedIn comments earned author replies. | sergebulaev/ | 4.3k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 210 | Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules. | davila7/ | 32k | 12 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 211 | Reverse-engineer what a company is building by scraping their job postings, careers page, LinkedIn Jobs, and engineering blog using TinyFish web agents. | tinyfish-io/ | 2.2k | — | ~3.6k | Automated safety check: Pass | MIT | 6 days ago |
| 212 | 212.Kol Engager Icp Find ICP-fit leads from KOL audiences on LinkedIn. An agent skill from gooseworks-ai/goose-skills. | gooseworks-ai/ | 1.2k | 1 repo | ~1.7k | Automated safety check: Notes | MIT | today |
| 213 | A skill your agent uses when a user provides DeepSeek web AI-search keywords, repeat count, target entity, and entity type, then needs repeated fresh-window crawls aggregated into JSON plus a Kami… | yaojingang/ | 868 | — | ~624 | Automated safety check: Pass | MIT | 7 days ago |
| 214 | 214.Browser Scrape 用 AutoCLI 二进制驱动用户已登录的 Chrome 抓取 Twitter/X、知乎、Bilibili、Reddit 等 55+ 站点 | ChanningLua/ | 273 | — | ~587 | Automated safety check: Notes | MIT | 27 days ago |
| 215 | Decision layer for researching social posts and trends that ranks discovery results and enforces a privacy screen before any identifying data leaves the local session. | kerpopule/ | 1k | — | ~2.5k | Automated safety check: Pass | MIT | yesterday |
| 216 | 216.Supercompress Compress bulky coding-agent context (tool dumps, logs, diffs, files, scrapes) before it burns tokens. | Supercompress/ | 106 | — | ~393 | Automated safety check: Pass | MIT | yesterday |
| 217 | Scrapes the Hacker News front page and returns the top 30 stories as JSON with rank, title, link, points and comment count. | garrytan/ | 136k | — | ~355 | Automated safety check: Pass | MIT | today |
| 218 | Extracts data from a web page through the Aside browser, using your real signed-in sessions, and returns one JSON document without changing anything. | garrytan/ | 136k | — | ~6.9k | Automated safety check: Notes | MIT | today |
| 219 | 219.Extract Run a quick extraction test against a URL or HTML file. An agent skill from RealEstateWebTools/property_web_scraper. | RealEstateWebTools/ | 132 | — | ~787 | Automated safety check: Pass | MIT | 2 mo ago |
| 220 | 220.Kuri Server Use kuri-server to automate Chrome via HTTP API — navigate pages, get a11y snapshots, interact with elements, capture network traffic (HAR), extract cookies, and bypass bot protection. | justrach/ | 365 | — | ~6.2k | Automated safety check: Notes | Unknown | 2 mo ago |
| 221 | Run, diagnose, and validate local recruitment crawler operations. | 849879772/ | 130 | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 222 | 222.Eval Loop Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. | malloydata/ | 116 | — | ~7.8k | Automated safety check: Pass | MIT | today |
| 223 | Extracts full detail data from a single Goofish (闲鱼/xianyu, goofish.com) second-hand item page. | browser-act/ | 6.1k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 224 | Scrapes second-hand item search results from Goofish (闲鱼/xianyu, goofish.com) — China's largest second-hand marketplace. | browser-act/ | 6.1k | 1 repo | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 225 | This skill helps users automatically scrape business data from Google Maps using the BrowserAct Google Maps API. | browser-act/ | 6.1k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 226 | Extracts business contact details from Google Maps search results and place detail pages, then visits each business website to collect emails, phone numbers, and social media profiles (Facebook… | browser-act/ | 6.1k | 1 repo | ~2.9k | Automated safety check: Pass | MIT | 1 mo ago |
| 227 | Search Taobao and Tmall product listings by keyword, returning paginated product cards with title, price, shop, image, sales, and tags. | browser-act/ | 6.1k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 228 | Browse a Taobao or Tmall shop's product catalog by shopId, returning paginated product listings with itemId and title. | browser-act/ | 6.1k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 229 | This skill helps users automatically extract complete Markdown content from any website via the BrowserAct Web Search Scraper API. | browser-act/ | 6.1k | 1 repo | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 230 | This skill helps users extract full article contents from WeChat using the BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 231 | This skill helps users automatically extract detailed video metrics and channel information from YouTube based on keyword searches using the BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 232 | This skill helps users automatically extract YouTube video transcripts and metadata in batch via the BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 233 | This skill helps users extract structured video list data and comment data from YouTube using the BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 234 | This skill helps users automatically extract YouTube video transcripts and metadata via the BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 235 | This skill helps users automatically extract channel-level and video detail data from a specific YouTube channel via BrowserAct API. | browser-act/ | 6.1k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 236 | Fix mypy --strict type errors in crawler files. An agent skill from opensanctions/opensanctions. | opensanctions/ | 832 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 237 | 237.Unifapi A skill your agent uses when working with UnifAPI public-data APIs or the UnifAPI MCP server: connecting OAuth MCP clients, discovering operations, calling social/search/scrape/news APIs… | unifapi-agent/ | 587 | — | ~741 | Automated safety check: Pass | MIT | 1 mo ago |
| 238 | A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages… | bgauryy/ | 949 | — | ~1.3k | Automated safety check: Pass | MIT | 5 days ago |
| 239 | A skill your agent uses when the user says "enrich this LinkedIn event", "enrich attendees of this event", "enrich this attendees CSV", "scrape and enrich LinkedIn event {URL}", "FullEnrich event… | Othmane-Khadri/ | 317 | — | ~1.4k | Automated safety check: Warn | MIT | 1 mo ago |
| 240 | Monitor competitor pricing, features, changelogs, dashboards, and product changes with Firecrawl. | firecrawl/ | 116 | — | ~604 | Automated safety check: Pass | ISC | yesterday |