A skill your agent uses when scraping web pages, automating browser interactions, crawling sites for content, or extracting data from rendered JavaScript pages

MITAuto-check passedData & Analytics

Install Web Crawling

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill web-crawling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace web-crawling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/productivity/cli-power-skills/skills/web-crawling .claude/skills/web-crawling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-crawling
GitHub stars
2.8k
Token cost
~1.5k tokens
SKILL.md length
311 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when scraping web pages, automating browser interactions, crawling sites for content, or extracting data from rendered JavaScript pages

  • Scraping web pages
  • SKILL.md covers When to Use, Tools, Patterns and Pipelines, plus 2 more sections
  • Calls node, python3 and duckdb
  • Automating browser interactions

What it does

Web Crawling is an agent skill from jeremylongshore/tons-of-skills-marketplace. Use when scraping web pages, automating browser interactions, crawling sites for content, or extracting data from rendered JavaScript pages

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. It works with JavaScript and Playwright. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Scraping web pages
  • Automating browser interactions
  • Crawling sites for content
  • Extracting data from rendered JavaScript pages

Example prompts

  • “/web-crawling”

Requirements

  • Python 3
  • Node.js
  • Pre-approved tools (allowed-tools): Bash(node*), Bash(python3*), Bash(scrapy*), Bash(katana*), Read, Write

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(node*)
    • Bash(python3*)
    • Bash(scrapy*)
    • Bash(katana*)
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • python3
    • duckdb

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Crawling loads about 1.5k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 311 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 311 words, ~1,534 tokens.

Download SKILL.mdSave it as .claude/skills/web-crawling/SKILL.md (or your agent's skills folder).
name
web-crawling
description
Use when scraping web pages, automating browser interactions, crawling sites for content, or extracting data from rendered JavaScript pages
allowed-tools
Bash(node*), Bash(python3*), Bash(scrapy*), Bash(katana*), Read, Write
version
1.0.0
author
ykotik
license
MIT

Web Crawling

When to Use

  • Scraping content from web pages (especially JavaScript-rendered)
  • Automating browser interactions (login, form submission, navigation)
  • Crawling a site to discover and extract structured data
  • Taking screenshots or generating PDFs of web pages
  • Discovering URLs and endpoints on a target site
  • Extracting data from sites with anti-bot protection

Tools

ToolPurposeWhen to choose
PlaywrightCross-browser automation (Chromium, Firefox, WebKit)JS-rendered pages, screenshots, complex interactions
PuppeteerChrome/Firefox automation via DevTools ProtocolChrome-specific automation, familiar API
ScrapyProduction-grade Python crawling frameworkLarge-scale structured crawling with pipelines
CrawleeNode.js crawling with anti-blocking built inSites with bot detection, proxy rotation needed
KatanaFast Go-based URL discovery crawlerEndpoint discovery, reconnaissance, speed

Patterns

Playwright: Scrape a page (inline Node.js script)
bash
node -e "
const { chromium } = require('playwright');
(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com');
  const title = await page.title();
  const text = await page.locator('body').innerText();
  console.log(JSON.stringify({ title, text }));
  await browser.close();
})();
"
Playwright: Take a screenshot
bash
node -e "
const { chromium } = require('playwright');
(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com');
  await page.screenshot({ path: 'screenshot.png', fullPage: true });
  await browser.close();
})();
"
Playwright: Extract structured data from a table
bash
node -e "
const { chromium } = require('playwright');
(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com/data');
  const rows = await page.locator('table tr').evaluateAll(trs =>
    trs.map(tr => Array.from(tr.querySelectorAll('td,th')).map(c => c.textContent.trim()))
  );
  console.log(JSON.stringify(rows));
  await browser.close();
})();
"
Playwright: Wait for JS content to load
bash
node -e "
const { chromium } = require('playwright');
(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com/spa');
  await page.waitForSelector('.content-loaded');
  const data = await page.locator('.content-loaded').innerText();
  console.log(data);
  await browser.close();
})();
"
Puppeteer: Generate PDF from a page
bash
node -e "
const puppeteer = require('puppeteer');
(async () => {
  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle0' });
  await page.pdf({ path: 'output.pdf', format: 'A4' });
  await browser.close();
})();
"
Katana: Discover all URLs on a site
bash
katana -u https://example.com -silent -d 3
Katana: Discover URLs and output as JSON
bash
katana -u https://example.com -silent -d 2 -jsonl
bash
katana -u https://example.com -headless -d 2 -silent
Scrapy: Quick single-page scrape (no project needed)
bash
scrapy fetch --nolog https://example.com | python3 -c "
import sys
from scrapy import Selector
html = sys.stdin.read()
sel = Selector(text=html)
for item in sel.css('h2::text').getall():
    print(item)
"
Crawlee: Basic crawler with anti-blocking
bash
node -e "
const { PlaywrightCrawler } = require('crawlee');
const crawler = new PlaywrightCrawler({
  async requestHandler({ page, request }) {
    const title = await page.title();
    console.log(JSON.stringify({ url: request.url, title }));
  },
  maxRequestsPerCrawl: 10,
});
crawler.run(['https://example.com']);
"

Pipelines

Discover URLs → filter interesting ones → scrape
bash
katana -u https://example.com -silent -d 2 | grep '/blog/' | head -5 | while read url; do
  node -e "
    const { chromium } = require('playwright');
    (async () => {
      const browser = await chromium.launch();
      const page = await browser.newPage();
      await page.goto('$url');
      const title = await page.title();
      const text = await page.locator('article').innerText().catch(() => '');
      console.log(JSON.stringify({ url: '$url', title, text: text.slice(0, 500) }));
      await browser.close();
    })();
  "
done

Each stage: Katana discovers URLs, grep filters to blog posts, Playwright scrapes each page.

Crawl → extract to JSON → query with DuckDB
bash
katana -u https://example.com -silent -d 2 -jsonl > crawl.jsonl
duckdb -c "SELECT endpoint, status_code, COUNT(*) FROM read_json_auto('crawl.jsonl') GROUP BY endpoint, status_code ORDER BY COUNT(*) DESC"

Each stage: Katana crawls to JSONL, DuckDB runs SQL on the crawl results.

Prefer Over

  • Prefer Playwright over curl for JavaScript-rendered pages — curl only gets raw HTML, Playwright executes JS
  • Prefer Katana over manual URL enumeration — fast, recursive, handles JS links
  • Prefer Scrapy over ad-hoc scripts for structured multi-page crawling — built-in rate limiting, pipelines, persistence

Do NOT Use When

  • Page content is available via a REST API — use the API directly (api-testing skill)
  • Extracting article text from a news URL — use newspaper4k (web-research skill), it's purpose-built
  • Simple static HTML that curl can fetch — use curl + jq/pup
  • The site explicitly prohibits scraping — respect robots.txt and ToS

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/productivity/cli-power-skills/skills/web-crawling of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Web Crawling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Crawling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Crawling this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: PassMIT
Anti Detect Browserantibrow/anti-detect-browser-skills17—~9.8kAutomated safety check: WarnMIT
CrawleeLeoYeAI/openclaw-master-skills2.2k—~4.7kAutomated safety check: PassMIT
Selenium Opinion Crawler123321kk/opinion-agent-ultimate107—~631Automated safety check: PassNone
Chrome Devtoolseinverne/dotfiles1211 repos~1.6kAutomated safety check: NotesApache-2.0
Agent Browseroxylabs/agent-skills875—~3kAutomated safety check: PassMIT

Similar skills

  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    17 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings
  • Crawlee

    LeoYeAI/openclaw-master-skills

    Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Selenium Opinion Crawler

    123321kk/opinion-agent-ultimate

    browser-based page capture and text extraction for public-opinion research.

    107 GitHub stars~631 tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check: notes
  • Agent Browser

    oxylabs/agent-skills

    Connects to Oxylabs remote agent browsers over the Chrome DevTools Protocol (CDP) with Playwright or Puppeteer.

    875 GitHub stars~3k tokensUpdated 10 days ago
    Productivity & AutomationAuto-check passed
  • Browser Automation

    alirezarezvani/claude-skills

    A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Web Crawling

What does Web Crawling do?

A skill your agent uses when scraping web pages, automating browser interactions, crawling sites for content, or extracting data from rendered JavaScript pages. Web Crawling is an agent skill from jeremylongshore/tons-of-skills-marketplace.

When should I use Web Crawling?

Web Crawling fits situations like: scraping web pages; automating browser interactions; crawling sites for content; extracting data from rendered JavaScript pages.

How do I install Web Crawling in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill web-crawling -a claude-code`. Or copy the skill folder (plugins/productivity/cli-power-skills/skills/web-crawling in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/web-crawling in your project. Claude Code loads it when a task matches its description.

How do I install Web Crawling in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill web-crawling -a codex`. Or copy the skill folder (plugins/productivity/cli-power-skills/skills/web-crawling in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/web-crawling in your project. Codex loads it when a task matches its description.

Can I use Web Crawling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill web-crawling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-crawling, .gemini/skills/web-crawling, .github/skills/web-crawling and .opencode/skills/web-crawling in your project.

What does Web Crawling need to run?

Going by SKILL.md and its folder, Web Crawling needs the command-line tools its instructions call (node, python3 and duckdb). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Bash(node*), Bash(python3*), Bash(scrapy*), Bash(katana*), Read, Write.

Does Web Crawling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Web Crawling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Crawling use?

Web Crawling is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Crawling use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Crawling?

Skills that share tags, products or a category with Web Crawling: Anti Detect Browser (antibrow/anti-detect-browser-skills, 17 stars), Crawlee (LeoYeAI/openclaw-master-skills, 2.2k stars), Selenium Opinion Crawler (123321kk/opinion-agent-ultimate, 107 stars) and Chrome Devtools (einverne/dotfiles, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Crawling?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.