Agent skill

Web Scrape

by 21pounder in 21pounder/terminalAgent

Intelligent web scraper with content extraction, multiple output formats, and error handling

Apache-2.0Auto-check passedData & Analytics

Install Web Scrape

skills CLI
$ npx skills add 21pounder/terminalAgent --skill web-scrape -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install 21pounder/terminalAgent web-scrape --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/21pounder/terminalAgent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/deepresearch/.claude/skills/web-scrape .claude/skills/web-scrape && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scrape
GitHub stars
120
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
369 words
Files
2 (incl. scripts)
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Intelligent web scraper with content extraction, multiple output formats, and error handling

  • Works in 6 steps: Navigate and Load → Capture Content → Close Browser → …
  • Tasks that involve Web scraping
  • SKILL.md covers Usage, Execution Flow, Smart Content Extraction and Output Formats, plus 5 more sections
  • Runs JavaScript scripts from its folder; reaches spa-app.com

What it does

Web Scrape is an agent skill from 21pounder/terminalAgent. Intelligent web scraper with content extraction, multiple output formats, and error handling

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/html_clean.js`).

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: A powerful terminal coding agent based on Claude Agent SDK. Better then Claude Code when start up a new project, or u have no idea. Just try it. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/web-scrape”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Navigate and Load
  2. Capture Content
  3. Close Browser
  4. Identify Content Type
  5. Filter Noise
  6. Structure the Content

What it can do on your machine

Read from SKILL.md and the folder at commit 989b029. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • spa-app.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scrape loads about 1.5k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 369 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from 21pounder/terminalAgent at commit 989b029, republished under its Apache-2.0 licence (© 21pounder). 369 words, ~1,478 tokens.

Download SKILL.mdSave it as .claude/skills/web-scrape/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
web-scrape
description
Intelligent web scraper with content extraction, multiple output formats, and error handling
version
3.0.0

Web Scraping Skill v3.0

Usage

/web-scrape <url> [options]

Options:

  • --format=markdown|json|text - Output format (default: markdown)
  • --full - Include full page content (skip smart extraction)
  • --screenshot - Also save a screenshot
  • --scroll - Scroll to load dynamic content (infinite scroll pages)

Examples:

/web-scrape https://example.com/article
/web-scrape https://news.site.com/story --format=json
/web-scrape https://spa-app.com/page --scroll --screenshot

Execution Flow

Phase 1: Navigate and Load
1. mcp__playwright__browser_navigate
   url: "<target URL>"

2. mcp__playwright__browser_wait_for
   time: 2  (allow initial render)

If --scroll option: Execute scroll sequence to trigger lazy loading:

3. mcp__playwright__browser_evaluate
   function: "async () => {
     for (let i = 0; i < 3; i++) {
       window.scrollTo(0, document.body.scrollHeight);
       await new Promise(r => setTimeout(r, 1000));
     }
     window.scrollTo(0, 0);
   }"
Phase 2: Capture Content
4. mcp__playwright__browser_snapshot
   → Returns full accessibility tree with all text content

If --screenshot option:

5. mcp__playwright__browser_take_screenshot
   filename: "scraped_<domain>_<timestamp>.png"
   fullPage: true
Phase 3: Close Browser
6. mcp__playwright__browser_close

Smart Content Extraction

After getting the snapshot, apply intelligent extraction:

Step 1: Identify Content Type
Page TypeIndicatorsExtraction Strategy
Article/Blog<article>, long paragraphs, date/authorExtract main article body
Product PagePrice, "Add to Cart", specsExtract title, price, description, specs
DocumentationCode blocks, headings hierarchyPreserve structure and code
List/SearchRepeated item patternsExtract as structured list
Landing PageHero section, CTAsExtract key messaging
Step 2: Filter Noise

ALWAYS REMOVE these elements from output:

  • Navigation menus and breadcrumbs
  • Footer content (copyright, links)
  • Sidebars (ads, related articles, social links)
  • Cookie banners and popups
  • Comments section (unless specifically requested)
  • Share buttons and social widgets
  • Login/signup prompts
Step 3: Structure the Content

For Articles:

markdown
# [Title]

**Source:** [URL]
**Date:** [if available]
**Author:** [if available]

---

[Main content in clean markdown]

For Product Pages:

markdown
# [Product Name]

**Price:** [price]
**Availability:** [in stock/out of stock]

## Description
[product description]

## Specifications
| Spec | Value |
|------|-------|
| ... | ... |

Output Formats

Markdown (default)

Clean, readable markdown with proper headings, lists, and formatting.

JSON
json
{
  "url": "https://...",
  "title": "Page Title",
  "type": "article|product|docs|list",
  "content": {
    "main": "...",
    "metadata": {}
  },
  "extracted_at": "ISO timestamp"
}
Text

Plain text with minimal formatting, suitable for further processing.


Error Handling

Show full SKILL.md (165 more words)Show less
Navigation Errors
ErrorDetectionAction
TimeoutPage doesn't load in 30sReport error, suggest retry
404 Not Found"404" in title/contentReport "Page not found"
403 Forbidden"403", "Access Denied"Report access restriction
CAPTCHA"captcha", "verify you're human"Report CAPTCHA detected, cannot proceed
Paywall"subscribe", "premium content"Extract visible content, note paywall
Recovery Actions
If page load fails:
1. Report the specific error to user
2. Suggest: "Try again?" or "Different URL?"
3. Close browser cleanly

If content is blocked:
1. Report what was detected (CAPTCHA/paywall/geo-block)
2. Extract any available preview content
3. Suggest alternatives if applicable

Advanced Scenarios

Single Page Applications (SPA)
1. Navigate to URL
2. Wait longer (3-5 seconds) for JS hydration
3. Use browser_wait_for with specific text if known
4. Then snapshot
Infinite Scroll Pages
1. Navigate
2. Execute scroll loop (see Phase 1)
3. Snapshot after scrolling completes
Pages with Click-to-Reveal Content
1. Snapshot first to identify clickable elements
2. Use browser_click on "Read more" / "Show all" buttons
3. Wait briefly
4. Snapshot again for full content
Multi-page Articles
1. Scrape first page
2. Identify "Next" or pagination links
3. Ask user: "Article has X pages. Scrape all?"
4. If yes, iterate through pages and combine

Performance Guidelines

MetricTargetHow
Speed< 15 secondsMinimal waits, parallel where possible
Token Usage< 5000 tokensSmart extraction, not full DOM
Reliability> 95% successProper error handling

Security Notes

  • Never execute arbitrary JavaScript from the page
  • Don't follow redirects to suspicious domains
  • Don't submit forms or click login buttons
  • Don't scrape pages that require authentication (unless user provides credentials flow)
  • Respect robots.txt when mentioned by user

Quick Reference

Minimum viable scrape (4 tool calls):

1. browser_navigate → 2. browser_wait_for → 3. browser_snapshot → 4. browser_close

Full-featured scrape (with scroll + screenshot):

1. browser_navigate
2. browser_wait_for
3. browser_evaluate (scroll)
4. browser_snapshot
5. browser_take_screenshot
6. browser_close

Remember: The goal is to deliver clean, useful content to the user, not raw HTML/DOM dumps.

© 21pounder, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in deepresearch/.claude/skills/web-scrape of 21pounder/terminalAgent.

  • SKILL.md
  • scripts/html_clean.js

Open the folder on GitHubat commit 989b029

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in 21pounder/terminalAgent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Web Scrape next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scrape compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scrape this skill21pounder/terminalAgent1201 repos~1.5kAutomated safety check: PassApache-2.0
Tmuxtrpc-group/trpc-agent-go1.8k24 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6971 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 24 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    697 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from 21pounder/terminalAgent

  • Deep Research

    21pounder/terminalAgent

    Conduct comprehensive deep research on any topic using Dify-powered workflow - searches documentation, academic papers, tutorials, APIs, best practices, and returns structured analysis with insights.

    120 GitHub starsUsed in 2 repos~883 tokens
    Auto-check: notes
  • Code Review

    21pounder/terminalAgent

    A skill your agent uses when user asks to "review code", "check for issues", "analyze code quality", "find bugs", or wants feedback on code implementation.

    120 GitHub starsUsed in 1 repo~650 tokens
    Auto-check passed
  • Git Commit

    21pounder/terminalAgent

    A skill your agent uses when user asks to "commit changes", "create a commit", "stage and commit", or wants help with git commit workflow.

    120 GitHub starsUsed in 1 repo~642 tokens
    Auto-check: notes

Questions about Web Scrape

What does Web Scrape do?

Intelligent web scraper with content extraction, multiple output formats, and error handling. Web Scrape is an agent skill from 21pounder/terminalAgent.

When should I use Web Scrape?

Web Scrape fits situations like: tasks that involve Web scraping.

How do I install Web Scrape in Claude Code?

Run `npx skills add 21pounder/terminalAgent --skill web-scrape -a claude-code`. Or copy the skill folder (deepresearch/.claude/skills/web-scrape in 21pounder/terminalAgent) into .claude/skills/web-scrape in your project. Claude Code loads it when a task matches its description.

How do I install Web Scrape in Codex?

Run `npx skills add 21pounder/terminalAgent --skill web-scrape -a codex`. Or copy the skill folder (deepresearch/.claude/skills/web-scrape in 21pounder/terminalAgent) into .agents/skills/web-scrape in your project. Codex loads it when a task matches its description.

Can I use Web Scrape in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add 21pounder/terminalAgent --skill web-scrape -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scrape, .gemini/skills/web-scrape, .github/skills/web-scrape and .opencode/skills/web-scrape in your project.

What does Web Scrape need to run?

Going by SKILL.md and its folder, Web Scrape needs JavaScript for the scripts in its folder. Our summary lists: Node.js.

Does Web Scrape access the network?

SKILL.md names 1 domain. In commands or code: spa-app.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Web Scrape safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Web Scrape use?

Web Scrape is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scrape use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Scrape?

Skills that share tags, products or a category with Web Scrape: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Ketch (1broseidon/ketch, 697 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scrape?

21pounder (a GitHub user) maintains it in 21pounder/terminalAgent, which has 120 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on May 28, 2026.

Source: 21pounder/terminalAgent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.