Agent skill

Web Scraping Automation

by aAAaqwq in aAAaqwq/AGI-Super-Team

自动化爬取网站数据和 API 接口。当用户需要抓取网页内容、调用 API、解析数据或创建爬虫脚本时使用此技能. An agent skill from aAAaqwq/AGI-Super-Team.

MITAuto-check: notesData & Analytics

Install Web Scraping Automation

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill web-scraping-automation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team web-scraping-automation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-scraping-automation .claude/skills/web-scraping-automation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scraping-automation
GitHub stars
105
Used in
1 other repo
Token cost
~795 tokens
SKILL.md length
115 words
Files
1
Skills in repo
167
Repo updated
First seen
Licence
MIT

At a glance

自动化爬取网站数据和 API 接口。当用户需要抓取网页内容、调用 API、解析数据或创建爬虫脚本时使用此技能. An agent skill from aAAaqwq/AGI-Super-Team.

  • Works in 3 steps: 简单网页爬取 → API 调用 → 动态网页爬取
  • Tasks that involve Web scraping
  • SKILL.md covers 功能说明, 使用场景, 技术栈 and 工作流程, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Web Scraping Automation is an agent skill from aAAaqwq/AGI-Super-Team. 自动化爬取网站数据和 API 接口。当用户需要抓取网页内容、调用 API、解析数据或创建爬虫脚本时使用此技能。

Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. It works with Selenium. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/web-scraping-automation”

Requirements

  • Python 3
  • A credential in YOUR_TOKEN
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit, WebFetch, WebSearch

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 简单网页爬取
  2. API 调用
  3. 动态网页爬取

What it can do on your machine

Read from SKILL.md and the folder at commit 331ecd3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scraping Automation loads about 795 tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 115 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~795

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit, WebFetch, WebSearch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 331ecd3, republished under its MIT licence (© aAAaqwq). 115 words, ~795 tokens.

Download SKILL.mdSave it as .claude/skills/web-scraping-automation/SKILL.md (or your agent's skills folder).
name
web-scraping-automation
description
自动化爬取网站数据和 API 接口。当用户需要抓取网页内容、调用 API、解析数据或创建爬虫脚本时使用此技能。
allowed-tools
Bash, Read, Write, Edit, WebFetch, WebSearch

网站爬取与 API 自动化

功能说明

此技能专门用于自动化网站数据爬取和 API 接口调用,包括:

  • 分析和爬取网站结构
  • 调用和测试 REST/GraphQL API
  • 创建自动化爬虫脚本
  • 数据解析和清洗
  • 处理反爬虫机制
  • 定时任务和数据存储

使用场景

  • "爬取这个网站的产品信息"
  • "帮我调用这个 API 并解析返回数据"
  • "创建一个脚本定时抓取新闻"
  • "分析这个网站的 API 接口文档"
  • "绕过这个网站的反爬虫限制"

技术栈

⚠️ 资源清理原则(强制)

所有涉及浏览器的爬取任务完成后,必须自动关闭 Chrome/Selenium 进程!

python
# Playwright 示例
from playwright.sync_api import sync_playwright

def scrape_website():
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        # ... 爬取逻辑 ...
        browser.close()

    # ⚠️ 强制清理残留进程
    import subprocess
    subprocess.run(['pkill', '-f', 'chrome'], capture_output=True)

# Selenium 示例
from selenium import webdriver

driver = webdriver.Chrome()
try:
    # ... 爬取逻辑 ...
    pass
finally:
    driver.quit()
    # ⚠️ 确保清理
    import subprocess
    subprocess.run(['pkill', '-f', 'chrome'], capture_output=True)

原因: 避免内存泄漏和资源占用,防止 Gateway CPU 100% 过载

Python 爬虫
  • requests:HTTP 请求库
  • BeautifulSoup4:HTML 解析
  • Scrapy:专业爬虫框架
  • Selenium:浏览器自动化
  • Playwright:现代浏览器自动化
JavaScript 爬虫
  • axios:HTTP 客户端
  • cheerio:服务端 jQuery
  • puppeteer:Chrome 自动化
  • node-fetch:Fetch API

工作流程

  1. 目标分析:

    • 检查网站结构和数据位置
    • 分析 API 接口和认证方式
    • 评估反爬虫机制
  2. 方案设计:

    • 选择合适的技术栈
    • 设计数据提取策略
    • 规划错误处理和重试机制
  3. 脚本开发:

    • 编写爬虫代码
    • 实现数据解析逻辑
    • 添加日志和监控
  4. 测试优化:

    • 验证数据准确性
    • 优化性能和稳定性
    • 处理边界情况

最佳实践

  • 遵守 robots.txt 规则
  • 设置合理的请求间隔
  • 使用 User-Agent 和请求头
  • 实现错误重试机制
  • 数据去重和验证
  • 使用代理池(如需要)
  • 保存原始数据和日志

常见场景示例

1. 简单网页爬取
python
import requests
from bs4 import BeautifulSoup

def scrape_website(url):
    headers = {'User-Agent': 'Mozilla/5.0'}
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')

    # 提取数据
    data = []
    for item in soup.select('.product'):
        data.append({
            'title': item.select_one('.title').text,
            'price': item.select_one('.price').text
        })
    return data
2. API 调用
python
import requests

def call_api(endpoint, params=None):
    headers = {
        'Authorization': 'Bearer YOUR_TOKEN',
        'Content-Type': 'application/json'
    }
    response = requests.get(endpoint, headers=headers, params=params)
    return response.json()
3. 动态网页爬取
python
from selenium import webdriver
from selenium.webdriver.common.by import By

def scrape_dynamic_page(url):
    driver = webdriver.Chrome()
    driver.get(url)

    # 等待页面加载
    driver.implicitly_wait(10)

    # 提取数据
    elements = driver.find_elements(By.CLASS_NAME, 'item')
    data = [elem.text for elem in elements]

    driver.quit()
    return data

反爬虫应对策略

  • 请求头伪装:模拟真实浏览器
  • 代理轮换:使用代理池
  • 验证码处理:OCR 或第三方服务
  • Cookie 管理:维护会话状态
  • 请求频率控制:避免触发限制
  • JavaScript 渲染:使用 Selenium/Playwright

数据存储方案

  • CSV/Excel:简单数据导出
  • JSON:结构化数据存储
  • 数据库:MySQL、PostgreSQL、MongoDB
  • 云存储:S3、OSS
  • 数据仓库:用于大规模数据分析

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/web-scraping-automation of aAAaqwq/AGI-Super-Team.

Open the folder on GitHubat commit 331ecd3

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Web Scraping Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scraping Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scraping Automation this skillaAAaqwq/AGI-Super-Team1051 repos~795Automated safety check: NotesMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Selenium Opinion Crawler123321kk/opinion-agent-ultimate107—~631Automated safety check: PassNone
Web Scraperrevfactory/harness-1001.3k—~1.7kAutomated safety check: PassApache-2.0
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7001 repos~3.9kAutomated safety check: PassMIT

Similar skills

  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Selenium Opinion Crawler

    123321kk/opinion-agent-ultimate

    browser-based page capture and text extraction for public-opinion research.

    107 GitHub stars~631 tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Web Scraper

    revfactory/harness-100

    Full pipeline for building web scraping systems with agent team collaboration.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    700 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 167 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Bankr Signals

    aAAaqwq/AGI-Super-Team

    Transaction-verified trading signals on Base blockchain. An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • Erc 8004

    aAAaqwq/AGI-Super-Team

    Register AI agents on Ethereum mainnet using ERC-8004 (Trustless Agents).

    105 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Web Scraping Automation

What does Web Scraping Automation do?

自动化爬取网站数据和 API 接口。当用户需要抓取网页内容、调用 API、解析数据或创建爬虫脚本时使用此技能. An agent skill from aAAaqwq/AGI-Super-Team. Web Scraping Automation is an agent skill from aAAaqwq/AGI-Super-Team.

When should I use Web Scraping Automation?

Web Scraping Automation fits situations like: tasks that involve Web scraping.

How do I install Web Scraping Automation in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill web-scraping-automation -a claude-code`. Or copy the skill folder (skills/web-scraping-automation in aAAaqwq/AGI-Super-Team) into .claude/skills/web-scraping-automation in your project. Claude Code loads it when a task matches its description.

How do I install Web Scraping Automation in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill web-scraping-automation -a codex`. Or copy the skill folder (skills/web-scraping-automation in aAAaqwq/AGI-Super-Team) into .agents/skills/web-scraping-automation in your project. Codex loads it when a task matches its description.

Can I use Web Scraping Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill web-scraping-automation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraping-automation, .gemini/skills/web-scraping-automation, .github/skills/web-scraping-automation and .opencode/skills/web-scraping-automation in your project.

What does Web Scraping Automation need to run?

SKILL.md names no scripts, command-line tools or credentials: Web Scraping Automation is instructions for the agent only. Our summary lists: Python 3; A credential in YOUR_TOKEN. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit, WebFetch, WebSearch.

Does Web Scraping Automation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Web Scraping Automation safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Web Scraping Automation use?

Web Scraping Automation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scraping Automation use?

About 795 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Scraping Automation?

Skills that share tags, products or a category with Web Scraping Automation: Python Executor (cortega26/chile-hub, 113 stars), Selenium Opinion Crawler (123321kk/opinion-agent-ultimate, 107 stars), Web Scraper (revfactory/harness-100, 1.3k stars) and Tmux (trpc-group/trpc-agent-go, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scraping Automation?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 167 skills in this directory. The repository was last updated on October 8, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.