Agent skill

Web Scraping

by cohen-liel in cohen-liel/hivemind

Web scraping and internet search patterns. An agent skill from cohen-liel/hivemind.

Apache-2.0Auto-check passedData & Analytics

Install Web Scraping

skills CLI
$ npx skills add cohen-liel/hivemind --skill web-scraping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cohen-liel/hivemind web-scraping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cohen-liel/hivemind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/web-scraping .claude/skills/web-scraping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scraping
GitHub stars
110
Token cost
~2k tokens
SKILL.md length
117 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Web scraping and internet search patterns. An agent skill from cohen-liel/hivemind.

  • Scraping websites
  • SKILL.md covers HTTP Scraping (httpx +…, Browser Automation (Playwright), Web Search (via SerpAPI or… and RSS Feed Reader, plus 4 more sections
  • Reaches serpapi.com and api.duckduckgo.com; needs SERPAPI_KEY
  • Extracting data from HTML

What it does

Web Scraping is an agent skill from cohen-liel/hivemind. Web scraping and internet search patterns. Use when scraping websites, crawling pages, extracting data from HTML, automating browser interactions, or fetching external web content programmatically.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: One prompt. A full AI engineering team. Go lie on the couch. 🧠. The licence is Apache-2.0.

When your agent uses it

  • Scraping websites
  • Extracting data from HTML
  • Automating browser interactions
  • Fetching external web content programmatically

Example prompts

  • “/web-scraping”

Requirements

  • Python 3
  • A credential in SERPAPI_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 918dd9b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • serpapi.com
    • api.duckduckgo.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SERPAPI_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scraping loads about 2k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 117 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cohen-liel/hivemind at commit 918dd9b, republished under its Apache-2.0 licence (© cohen-liel). 117 words, ~1,981 tokens.

Download SKILL.mdSave it as .claude/skills/web-scraping/SKILL.md (or your agent's skills folder).
name
web-scraping
description
Web scraping and internet search patterns. Use when scraping websites, crawling pages, extracting data from HTML, automating browser interactions, or fetching external web content programmatically.

Web Scraping & Internet Search Patterns

HTTP Scraping (httpx + BeautifulSoup)

python
import httpx
from bs4 import BeautifulSoup
import asyncio

async def scrape_page(url: str) -> dict:
    """Fetch and parse a single page."""
    headers = {
        "User-Agent": "Mozilla/5.0 (compatible; MyBot/1.0; +https://example.com/bot)"
    }
    async with httpx.AsyncClient(headers=headers, follow_redirects=True, timeout=30) as client:
        response = await client.get(url)
        response.raise_for_status()

    soup = BeautifulSoup(response.text, "lxml")

    return {
        "title": soup.find("title").get_text(strip=True) if soup.find("title") else "",
        "headings": [h.get_text(strip=True) for h in soup.find_all(["h1", "h2", "h3"])],
        "links": [a["href"] for a in soup.find_all("a", href=True)],
        "text": soup.get_text(separator="\n", strip=True)[:5000],
    }

async def scrape_many(urls: list[str], max_concurrent: int = 5) -> list[dict]:
    """Scrape many URLs with concurrency limit."""
    sem = asyncio.Semaphore(max_concurrent)

    async def fetch_one(url):
        async with sem:
            try:
                return await scrape_page(url)
            except Exception as e:
                return {"url": url, "error": str(e)}

    return await asyncio.gather(*[fetch_one(url) for url in urls])

Browser Automation (Playwright)

python
from playwright.async_api import async_playwright

async def scrape_with_browser(url: str) -> str:
    """Use for JS-heavy sites that need a real browser."""
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()

        # Block images/fonts for speed
        await page.route("**/*.{png,jpg,gif,webp,svg,woff,woff2}", lambda r: r.abort())

        await page.goto(url, wait_until="networkidle", timeout=30000)

        # Wait for specific element
        await page.wait_for_selector(".content", timeout=10000)

        # Extract data
        text = await page.inner_text("body")
        links = await page.eval_on_selector_all("a[href]", "els => els.map(e => e.href)")

        await browser.close()
        return text

async def fill_form_and_submit(url: str, form_data: dict) -> str:
    """Automate form submission."""
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto(url)

        for selector, value in form_data.items():
            await page.fill(selector, value)

        await page.click("button[type=submit]")
        await page.wait_for_load_state("networkidle")
        result = await page.inner_text("body")
        await browser.close()
        return result

Web Search (via SerpAPI or DuckDuckGo)

python
import httpx

async def search_web(query: str, num_results: int = 10) -> list[dict]:
    """Search the web using SerpAPI."""
    async with httpx.AsyncClient() as client:
        response = await client.get(
            "https://serpapi.com/search",
            params={
                "q": query,
                "num": num_results,
                "api_key": settings.SERPAPI_KEY,
                "engine": "google",
            },
            timeout=30,
        )
        data = response.json()

    return [
        {
            "title": r.get("title"),
            "url": r.get("link"),
            "snippet": r.get("snippet"),
        }
        for r in data.get("organic_results", [])
    ]

async def search_duckduckgo(query: str) -> list[dict]:
    """Free alternative — DuckDuckGo instant answers (no API key)."""
    async with httpx.AsyncClient() as client:
        response = await client.get(
            "https://api.duckduckgo.com/",
            params={"q": query, "format": "json", "no_html": 1},
            timeout=15,
        )
    data = response.json()
    results = []
    for r in data.get("Results", []):
        results.append({"title": r["Text"], "url": r["FirstURL"]})
    return results

RSS Feed Reader

python
import feedparser
import httpx

async def read_rss(feed_url: str) -> list[dict]:
    """Parse RSS/Atom feed."""
    async with httpx.AsyncClient() as client:
        response = await client.get(feed_url, timeout=15)
    feed = feedparser.parse(response.text)
    return [
        {
            "title": entry.title,
            "url": entry.link,
            "summary": entry.get("summary", ""),
            "published": entry.get("published", ""),
        }
        for entry in feed.entries[:20]
    ]

Data Extraction (Structured)

python
from pydantic import BaseModel
from typing import Optional
import re

class ProductData(BaseModel):
    name: str
    price: Optional[float]
    description: str
    image_url: Optional[str]

def extract_product(soup: BeautifulSoup, url: str) -> ProductData:
    """Extract structured product data."""
    # Try JSON-LD schema first (most reliable)
    json_ld = soup.find("script", type="application/ld+json")
    if json_ld:
        import json
        data = json.loads(json_ld.string)
        if data.get("@type") == "Product":
            return ProductData(
                name=data["name"],
                price=float(data.get("offers", {}).get("price", 0)),
                description=data.get("description", ""),
                image_url=data.get("image"),
            )

    # Fallback: heuristic selectors
    name = soup.find(["h1", '[class*="title"]', '[itemprop="name"]'])
    price_el = soup.find(['[class*="price"]', '[itemprop="price"]'])
    price_text = price_el.get_text() if price_el else ""
    price = float(re.search(r'[\d.]+', price_text).group()) if re.search(r'[\d.]+', price_text) else None

    return ProductData(
        name=name.get_text(strip=True) if name else "",
        price=price,
        description="",
        image_url=None,
    )

Politeness & Rate Limiting

python
import time, random

class PoliteScraper:
    def __init__(self, delay_range=(1, 3)):
        self.delay_range = delay_range
        self.last_request = {}

    async def fetch(self, url: str, session: httpx.AsyncClient) -> str:
        domain = httpx.URL(url).host

        # Respect per-domain delay
        if domain in self.last_request:
            elapsed = time.time() - self.last_request[domain]
            min_delay = random.uniform(*self.delay_range)
            if elapsed < min_delay:
                await asyncio.sleep(min_delay - elapsed)

        response = await session.get(url)
        self.last_request[domain] = time.time()
        return response.text

Robots.txt Compliance

python
from urllib.robotparser import RobotFileParser
from urllib.parse import urljoin

def can_scrape(url: str, user_agent: str = "*") -> bool:
    """Check robots.txt before scraping."""
    parsed = httpx.URL(url)
    robots_url = f"{parsed.scheme}://{parsed.host}/robots.txt"
    rp = RobotFileParser()
    rp.set_url(robots_url)
    rp.read()
    return rp.can_fetch(user_agent, url)

Rules

  • ALWAYS check robots.txt before scraping — respect disallow rules
  • Set a descriptive User-Agent (not a fake browser agent) for bots
  • Rate limit: minimum 1-2 second delay between requests to same domain
  • Use Playwright only when necessary (JS rendering) — httpx is 10x faster
  • Cache scraped pages (Redis/disk) to avoid re-fetching during development
  • Handle rate limit responses (429): exponential backoff with jitter
  • Never scrape personal data without legal basis (GDPR compliance)
  • For large crawls: use Scrapy framework instead of rolling your own
  • Store raw HTML alongside extracted data for re-processing

© cohen-liel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/web-scraping of cohen-liel/hivemind.

Open the folder on GitHubat commit 918dd9b

Compare with similar skills

Web Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scraping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scraping this skillcohen-liel/hivemind110—~2kAutomated safety check: PassApache-2.0
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7021 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    702 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 11 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from cohen-liel/hivemind

All 33 skills in this repo
  • API Design

    cohen-liel/hivemind

    REST API design principles and best practices. An agent skill from cohen-liel/hivemind.

    110 GitHub stars~857 tokensUpdated 5 mo ago
    Auto-check passed
  • Async Python

    cohen-liel/hivemind

    Python asyncio patterns for high-performance async code. An agent skill from cohen-liel/hivemind.

    110 GitHub stars~992 tokensUpdated 5 mo ago
    Auto-check passed
  • Celery Tasks

    cohen-liel/hivemind

    Celery background task patterns for Python apps. An agent skill from cohen-liel/hivemind.

    110 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Docker Deployment

    cohen-liel/hivemind

    Docker, docker-compose, and deployment configuration best practices.

    110 GitHub stars~674 tokensUpdated 5 mo ago
    Auto-check passed
  • E2E Testing

    cohen-liel/hivemind

    End-to-end testing patterns with Playwright. An agent skill from cohen-liel/hivemind.

    110 GitHub stars~1.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Email Service

    cohen-liel/hivemind

    Email sending patterns for transactional and marketing emails.

    110 GitHub stars~1.8k tokensUpdated 5 mo ago
    Auto-check passed

Questions about Web Scraping

What does Web Scraping do?

Web scraping and internet search patterns. An agent skill from cohen-liel/hivemind. Web Scraping is an agent skill from cohen-liel/hivemind. Web scraping and internet search patterns.

When should I use Web Scraping?

Web Scraping fits situations like: scraping websites; extracting data from HTML; automating browser interactions; fetching external web content programmatically.

How do I install Web Scraping in Claude Code?

Run `npx skills add cohen-liel/hivemind --skill web-scraping -a claude-code`. Or copy the skill folder (.claude/skills/web-scraping in cohen-liel/hivemind) into .claude/skills/web-scraping in your project. Claude Code loads it when a task matches its description.

How do I install Web Scraping in Codex?

Run `npx skills add cohen-liel/hivemind --skill web-scraping -a codex`. Or copy the skill folder (.claude/skills/web-scraping in cohen-liel/hivemind) into .agents/skills/web-scraping in your project. Codex loads it when a task matches its description.

Can I use Web Scraping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cohen-liel/hivemind --skill web-scraping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraping, .gemini/skills/web-scraping, .github/skills/web-scraping and .opencode/skills/web-scraping in your project.

What does Web Scraping need to run?

Going by SKILL.md and its folder, Web Scraping needs credentials named SERPAPI_KEY. Our summary lists: Python 3; A credential in SERPAPI_KEY.

Does Web Scraping access the network?

SKILL.md names 2 domains. In commands or code: serpapi.com and api.duckduckgo.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web Scraping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Scraping use?

Web Scraping is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scraping use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Scraping?

Skills that share tags, products or a category with Web Scraping: Tmux (trpc-group/trpc-agent-go, 1.9k stars), Ketch (1broseidon/ketch, 702 stars), Crawl4AI Web Scraping (smallnest/goclaw, 599 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scraping?

cohen-liel (a GitHub user) maintains it in cohen-liel/hivemind, which has 110 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on April 18, 2026.

Source: cohen-liel/hivemind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.