Agent skill

Web Scraper

by shobcoder in shobcoder/shob

Scrape, crawl, and extract data from websites. An agent skill from shobcoder/shob.

MITAuto-check passedData & Analytics

Install Web Scraper

skills CLI
$ npx skills add shobcoder/shob --skill web-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shobcoder/shob web-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shobcoder/shob.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-scraper .claude/skills/web-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scraper
GitHub stars
577
Token cost
~700 tokens
SKILL.md length
171 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Scrape, crawl, and extract data from websites. An agent skill from shobcoder/shob.

  • Works in 5 steps: Start with simpler extraction before… → Use specific prompts for targeted data → Respect website terms of service → …
  • Users ask to scrape web pages
  • SKILL.md covers Overview, When to Use, Tools Available and Usage Patterns, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Web Scraper is an agent skill from shobcoder/shob. Scrape, crawl, and extract data from websites. Use when users ask to scrape web pages, extract content, crawl websites, or collect data from the internet.

Its SKILL.md is about 700 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: Shob – an AI agent that delivers high-quality coding & automation work. The licence is MIT.

When your agent uses it

  • Users ask to scrape web pages
  • Extract content
  • Collect data from the internet

Example prompts

  • “/web-scraper”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Start with simpler extraction before complex patterns
  2. Use specific prompts for targeted data
  3. Respect website terms of service
  4. Add delays between requests when scraping multiple pages
  5. Handle errors gracefully with try/catch

What it can do on your machine

Read from SKILL.md and the folder at commit 14831ba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scraper loads about 700 tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 171 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~700

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shobcoder/shob at commit 14831ba, republished under its MIT licence (© shobcoder). 171 words, ~700 tokens.

Download SKILL.mdSave it as .claude/skills/web-scraper/SKILL.md (or your agent's skills folder).
name
web-scraper
description
Scrape, crawl, and extract data from websites. Use when users ask to scrape web pages, extract content, crawl websites, or collect data from the internet.

Web Scraper

Overview

Extract content and data from websites using various techniques including crawling, scraping, and structured data extraction.

When to Use

  • Extract text content from web pages
  • Crawl entire websites
  • Collect structured data
  • Research and gather information
  • Monitor website changes
  • Extract tables and lists

Tools Available

Content Extraction
javascript
// Use extract_content_from_websites for structured extraction
// Supports batch processing of multiple URLs
// Returns JSON format with extracted content
Task Format
javascript
{
    tasks: [
        {
            url: "https://example.com",
            prompt: "Extract specific information",
            task_name: "optional_name"
        }
    ]
}

Usage Patterns

Simple Content Extraction
javascript
// Extract main content from a page
const result = await extract_content_from_websites({
    tasks: [{
        url: "https://news.example.com/article",
        prompt: "Extract the title, author, date, and main content"
    }]
});
Batch URL Processing
javascript
// Process multiple URLs in parallel
const urls = [
    "https://site.com/page1",
    "https://site.com/page2",
    "https://site.com/page3"
];

const results = await extract_content_from_websites({
    tasks: urls.map((url, i) => ({
        url,
        prompt: "Extract all product information, prices, and descriptions",
        task_name: `product_${i}`
    }))
});
Data Mining
javascript
// Extract structured data like prices, reviews, specifications
const data = await extract_content_from_websites({
    tasks: [{
        url: "https://ecommerce.example.com/products",
        prompt: "Extract product name, price, rating, and availability for all products listed"
    }]
});

Extraction Modes

Auto Mode (Default)
  • Attempts HTTP GET first
  • Falls back to browser rendering for CSR pages
  • Best for most websites
Curl Only Mode
  • Fast direct HTTP requests
  • Best for static HTML pages
  • May fail on JavaScript-heavy sites
Browser Only Mode
  • Full browser rendering
  • Handles dynamic content
  • Slower but more comprehensive

Best Practices

  1. Start with simpler extraction before complex patterns
  2. Use specific prompts for targeted data
  3. Respect website terms of service
  4. Add delays between requests when scraping multiple pages
  5. Handle errors gracefully with try/catch

Data Handling

  • Returns JSON format for easy processing
  • Handles batch operations efficiently
  • Supports pagination when needed
  • Maintains data structure in results

© shobcoder, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/web-scraper of shobcoder/shob.

Open the folder on GitHubat commit 14831ba

Compare with similar skills

Web Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scraper this skillshobcoder/shob577—~700Automated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.8k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6961 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    696 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from shobcoder/shob

  • Deep Research

    shobcoder/shob

    GOD MODE deep research skill. An agent skill from shobcoder/shob.

    577 GitHub stars~2.1k tokensUpdated 19 days ago
    Auto-check passed
  • Deep Research Agent

    shobcoder/shob

    Comprehensive research agent for in-depth investigation. An agent skill from shobcoder/shob.

    577 GitHub stars~1.9k tokensUpdated 19 days ago
    Auto-check passed
  • Effect

    shobcoder/shob

    Work with Effect v4 / effect-smol TypeScript code in this repo

    577 GitHub starsUsed in 5 repos~694 tokens
    Auto-check passed
  • Memory

    shobcoder/shob

    Persistent, token-efficient project memory. An agent skill from shobcoder/shob.

    577 GitHub stars~2.5k tokensUpdated 19 days ago
    Auto-check passed
  • UI UX Pro Max

    shobcoder/shob

    UI/UX design intelligence expert for web and mobile applications.

    577 GitHub stars~3.7k tokensUpdated 19 days ago
    Auto-check passed

Questions about Web Scraper

What does Web Scraper do?

Scrape, crawl, and extract data from websites. An agent skill from shobcoder/shob. Web Scraper is an agent skill from shobcoder/shob. Scrape, crawl, and extract data from websites.

When should I use Web Scraper?

Web Scraper fits situations like: users ask to scrape web pages; extract content; collect data from the internet.

How do I install Web Scraper in Claude Code?

Run `npx skills add shobcoder/shob --skill web-scraper -a claude-code`. Or copy the skill folder (skills/web-scraper in shobcoder/shob) into .claude/skills/web-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Web Scraper in Codex?

Run `npx skills add shobcoder/shob --skill web-scraper -a codex`. Or copy the skill folder (skills/web-scraper in shobcoder/shob) into .agents/skills/web-scraper in your project. Codex loads it when a task matches its description.

Can I use Web Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shobcoder/shob --skill web-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraper, .gemini/skills/web-scraper, .github/skills/web-scraper and .opencode/skills/web-scraper in your project.

What does Web Scraper need to run?

SKILL.md names no scripts, command-line tools or credentials: Web Scraper is instructions for the agent only.

Does Web Scraper access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Web Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Scraper use?

Web Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scraper use?

About 700 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Scraper?

Skills that share tags, products or a category with Web Scraper: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Ketch (1broseidon/ketch, 696 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scraper?

shobcoder (a GitHub organization) maintains it in shobcoder/shob, which has 577 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 18, 2026.

Source: shobcoder/shob on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.