Agent skill

Web Scraping

by ginlix-ai in ginlix-ai/LangAlpha

Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…

MITAuto-check passedData & Analytics

Install Web Scraping

skills CLI
$ npx skills add ginlix-ai/LangAlpha --skill web-scraping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ginlix-ai/LangAlpha web-scraping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/alternative_data/skills/web-scraping .claude/skills/web-scraping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scraping
GitHub stars
1.8k
Token cost
~2.1k tokens
SKILL.md length
491 words
Files
2 (incl. references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…

  • Works in 2 steps: MCP tools (scrape_page, scrape_pages) —… → Direct Scrapling Python API — for…
  • Tasks that involve Web scraping
  • SKILL.md covers Overview, MCP Tools, Direct Python API (Advanced) and Converting HTML to Markdown, plus 1 more section
  • Reaches protected-site.com and spa-site.com

What it does

Web Scraping is an agent skill from ginlix-ai/LangAlpha. Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors, sessions, and spiders

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api-reference.md`).

It sits in Data & Analytics, covering Web scraping. It works with Python. The repository describes itself as: Claude Code for Financial Market. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/web-scraping”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. MCP tools (scrape_page, scrape_pages) — recommended for straight "give me this page's content". Synchronous, return dicts.
  2. Direct Scrapling Python API — for CSS/XPath selectors, sessions, logins, and multi-page spiders. Async, returns Page objects with .css() /…

What it can do on your machine

Read from SKILL.md and the folder at commit 2855e43. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • protected-site.com
    • spa-site.com
    • spa-website.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scraping loads about 2.1k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ginlix-ai/LangAlpha at commit 2855e43, republished under its MIT licence (© ginlix-ai). 491 words, ~2,120 tokens.

Download SKILL.mdSave it as .claude/skills/web-scraping/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
web-scraping
description
Web scraping: scrape_page / scrape_pages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors, sessions, and spiders
license
MIT

Web Scraping

Overview

Two ways to scrape in the sandbox:

  1. MCP tools (scrape_page, scrape_pages) — recommended for straight "give me this page's content". Synchronous, return dicts.
  2. Direct Scrapling Python API — for CSS/XPath selectors, sessions, logins, and multi-page spiders. Async, returns Page objects with .css() / .xpath().

Quick fetches can run inline via ExecuteCode. For spiders, multi-URL crawls, or anything you'll iterate on, write the scraper to <task_name>/scraper.py and run it via Bash — edit-and-rerun beats resubmitting code.

MCP Tools

Import from tools.scrape. Synchronous — no await.

python
from tools.scrape import scrape_page, scrape_pages
Signatures
python
scrape_page(url: str, mode: str = "fast", extraction: str = "markdown",
            timeout: float = 30.0, solve_cloudflare: bool = False) -> dict

scrape_pages(urls: list[str], mode: str = "fast", extraction: str = "markdown",
             timeout: float = 30.0, solve_cloudflare: bool = False) -> dict
Parameters
ParamDefaultNotes
mode"fast""fast" plain HTTP · "browser" JS rendering · "stealth" bot-protected sites
extraction"markdown""markdown" (article text, cleaned) · "html" (raw) · "text" (plain)
timeout30.0Per-fetch seconds, 1–60 — seconds in every mode, not ms
solve_cloudflareFalseOnly meaningful with mode="stealth"
urls—scrape_pages only; max 10 per call

Escalate modes only as needed: start fast, go to browser when the page needs JavaScript, stealth when you're getting blocked, and add solve_cloudflare=True only if stealth still returns a challenge page.

Return shape

scrape_page returns a flat dict:

python
{
    "url": "https://example.com",
    "status": 200,
    "title": "Example Domain",
    "content": "# Example Domain\n\nThis domain is for use in...",  # str
    "extraction": "markdown",
    "mode": "fast",
}
  • content is a plain string, not a list — use it directly, never content[0] (that yields a single character).
  • content is truncated to 400,000 chars.
  • No .css() / .xpath() / .body / .headers / .cookies — for selectors use the direct Python API below, or parse extraction="html" with BeautifulSoup.

scrape_pages wraps them:

python
{
    "results": [ ... ],  # one entry per input URL, in input order
    "count": 3,
}
Errors

Errors are returned, never raised. Always check for "error" before reading content.

python
res = scrape_page(url="https://example.com")
if "error" in res:
    print(res["error"], res["detail"])
else:
    print(res["content"])

Per-URL errors — appear as {"error", "detail", "url"} entries inside scrape_pages["results"], or as the whole return of scrape_page:

CodeMeaning
invalid_urlNot an http:// / https:// URL
fetch_failedNetwork, DNS, timeout, or browser failure
extract_failedPage fetched but the extractor failed on the markup; the entry still carries status
scrape_failedUnexpected internal failure for that one URL

Whole-call errors — the entire return is {"error", "detail"}, no results:

CodeMeaning
invalid_mode / invalid_extraction / invalid_timeoutBad argument value
invalid_urlsscrape_pages got an empty list or more than 10 URLs

One bad URL never sinks a batch. scrape_pages always returns one entry per input URL, in input order — failures come back as error entries alongside the successes.

Show full SKILL.md (149 more words)Show less
Examples
python
from tools.scrape import scrape_page, scrape_pages

# Single page → markdown
res = scrape_page(url="https://example.com")
if "error" not in res:
    print(res["title"], res["status"], len(res["content"]))

# JS-rendered page
res = scrape_page(url="https://spa-site.com", mode="browser", timeout=60)

# Bot-protected page
res = scrape_page(url="https://protected-site.com", mode="stealth", solve_cloudflare=True)

# Batch — split successes from failures
batch = scrape_pages(urls=[...], mode="fast")   # <= 10 URLs
pages = [r for r in batch["results"] if "error" not in r]
failed = [(r["url"], r["error"]) for r in batch["results"] if "error" in r]

# Raw HTML when you need to parse structure yourself
res = scrape_page(url="https://example.com", extraction="html")
from bs4 import BeautifulSoup
soup = BeautifulSoup(res["content"], "html.parser")
titles = [h1.get_text() for h1 in soup.find_all("h1")]

Batches run concurrently — 8 at a time in fast mode, 2 at a time in browser / stealth (browser sessions are memory-heavy). More than 10 URLs means more than one call.


Direct Python API (Advanced)

For selectors, sessions, spiders, or when you need the full Page object. Requires imports. Async.

Fetcher (Fast HTTP — Tier 1)
python
from scrapling.fetchers import AsyncFetcher

page = await AsyncFetcher.get("https://example.com", stealthy_headers=True)
print(page.status)       # 200
print(page.body)         # Raw bytes
print(page.headers)      # Response headers

# CSS selectors (Scrapy-style pseudo-elements)
titles = page.css("h1::text").getall()
links = page.css("a::attr(href)").getall()

# XPath
items = page.xpath("//div[@class='item']/text()").getall()

# BeautifulSoup-style
divs = page.find_all("div", class_="content")
DynamicFetcher (Browser — Tier 2)
python
from scrapling.fetchers import DynamicFetcher

page = await DynamicFetcher.async_fetch(
    "https://spa-website.com",
    headless=True,
    network_idle=True,
    disable_resources=True,
    timeout=30000,          # milliseconds here, unlike the MCP tools
    wait_selector=".data-table",
)
rows = page.css("table.data-table tr")
for row in rows:
    cells = row.css("td::text").getall()
StealthyFetcher (Anti-Bot — Tier 3)
python
from scrapling.fetchers import StealthyFetcher

page = await StealthyFetcher.async_fetch(
    "https://protected-site.com",
    headless=True,
    solve_cloudflare=True,
    network_idle=True,
)
Sessions (Persistent Connections)
python
from scrapling.fetchers import FetcherSession

with FetcherSession(impersonate="chrome") as session:
    login_page = session.post("https://site.com/login", data={...})
    dashboard = session.get("https://site.com/dashboard")
    data = dashboard.css(".user-data::text").getall()
Spider (Multi-Page Crawl)
python
from scrapling.spiders import Spider, Request, Response

class PriceScraper(Spider):
    name = "prices"
    start_urls = ["https://example.com/products"]
    concurrent_requests = 5

    async def parse(self, response: Response):
        for product in response.css(".product"):
            yield {
                "name": product.css(".name::text").get(),
                "price": product.css(".price::text").get(),
            }
        next_page = response.css("a.next::attr(href)").get()
        if next_page:
            yield Request(next_page)

spider = PriceScraper()
result = spider.start()
result.items.to_json("<task_name>/data/prices.json")

Converting HTML to Markdown

Only needed when you fetched HTML yourself — scrape_page(extraction="markdown") already does this.

python
import html_to_markdown

markdown = html_to_markdown.convert(
    html_string, html_to_markdown.ConversionOptions(extract_metadata=False)
).content

# Article-only extraction (strips nav/ads/boilerplate)
import trafilatura

article = trafilatura.extract(html_string, output_format="markdown", favor_recall=True)

When to Use Which

NeedUse
Quick page content as markdownscrape_page()
Several known URLs at oncescrape_pages() (≤10 per call)
Extract specific elements (CSS/XPath)Direct Python API with selectors
Login + scrape authenticated pagesDirect Python API with sessions
Crawl many pages with paginationDirect Python API with Spider
Bypass Cloudflarescrape_page(mode="stealth", solve_cloudflare=True) or direct StealthyFetcher
Save results to fileDirect Python API (spider .to_json())

© ginlix-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/alternative_data/skills/web-scraping of ginlix-ai/LangAlpha.

  • SKILL.md
  • references/api-reference.md

Open the folder on GitHubat commit 2855e43

Compare with similar skills

Web Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scraping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scraping this skillginlix-ai/LangAlpha1.8k—~2.1kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT
Google Maps ScraperMahanaicoach/google-maps-scraper-kit1.3k—~2.8kAutomated safety check: PassMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 11 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Google Maps Scraper

    Mahanaicoach/google-maps-scraper-kit

    Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.

    1.3k GitHub stars~2.8k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Scrapling

    Cedriccmh/claude-code-skill-scrapling

    使用 scrapling 进行网页抓取和数据提取。根据目标网站特征自动选择最佳 Fetcher, 生成并执行 Python 脚本完成任务。Use when: (1) 抓取/爬取网页内容或数据(scrape, crawl, fetch page, extract data) (2) 需要绕过 Cloudflare/WAF 等反爬保护 (3) 登录后抓取受保护页面 (4) 解析已有 HTML…

    446 GitHub stars~1.1k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from ginlix-ai/LangAlpha

All 37 skills in this repo
  • Investment Deck Check

    ginlix-ai/LangAlpha

    Quality-checks an investment deck in .pptx form before it goes out: number consistency, chart and narrative alignment, source coverage, language and a circulation verdict.

    1.8k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Equity Initiation Report

    ginlix-ai/LangAlpha

    Produces a first-time equity research initiation report in five tasks: company research, financial model, valuation, charts and a DOCX report.

    1.8k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Builds or repairs an integrated income statement, balance sheet and cash flow model in Excel with live formulas, supporting schedules, scenarios and a Checks sheet.

    1.8k GitHub stars~5.4k tokensUpdated today
    Auto-check passed
  • Financial Model Checker

    ginlix-ai/LangAlpha

    Audits an existing Excel financial model without editing it, checking structure, formulas, integrity identities and source tie-out, and ends in a prioritized issue log.

    1.8k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • DCF Model Builder

    ginlix-ai/LangAlpha

    Builds a live Excel DCF valuation workbook with free cash flow projections, WACC, terminal value, three scenarios, sensitivity grids and a reverse DCF.

    1.8k GitHub stars~7.7k tokensUpdated today
    Auto-check passed
  • Builds Word files with python-docx, edits existing ones in place with tracked changes and comments, then renders and validates the result.

    1.8k GitHub stars~4.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Web Scraping

What does Web Scraping do?

Web scraping: scrapepage / scrapepages MCP tools for fetching pages as markdown, HTML, or text (fast HTTP, browser rendering, anti-bot stealth), plus the direct Scrapling Python API for selectors…. Web Scraping is an agent skill from ginlix-ai/LangAlpha.

When should I use Web Scraping?

Web Scraping fits situations like: tasks that involve Web scraping.

How do I install Web Scraping in Claude Code?

Run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a claude-code`. Or copy the skill folder (plugins/alternative_data/skills/web-scraping in ginlix-ai/LangAlpha) into .claude/skills/web-scraping in your project. Claude Code loads it when a task matches its description.

How do I install Web Scraping in Codex?

Run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a codex`. Or copy the skill folder (plugins/alternative_data/skills/web-scraping in ginlix-ai/LangAlpha) into .agents/skills/web-scraping in your project. Codex loads it when a task matches its description.

Can I use Web Scraping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ginlix-ai/LangAlpha --skill web-scraping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraping, .gemini/skills/web-scraping, .github/skills/web-scraping and .opencode/skills/web-scraping in your project.

What does Web Scraping need to run?

SKILL.md names no scripts, command-line tools or credentials: Web Scraping is instructions for the agent only. Our summary lists: Python 3.

Does Web Scraping access the network?

SKILL.md names 3 domains. In commands or code: protected-site.com, spa-site.com and spa-website.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web Scraping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Scraping use?

Web Scraping is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scraping use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Web Scraping?

Skills that share tags, products or a category with Web Scraping: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars), Ax (yusukebe/ax, 719 stars) and Google Maps Scraper (Mahanaicoach/google-maps-scraper-kit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scraping?

ginlix-ai (a GitHub organization) maintains it in ginlix-ai/LangAlpha, which has 1,811 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 10, 2026.

Source: ginlix-ai/LangAlpha on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.