Agent skill

Easy Spider Guide

by wentorai in wentorai/research-plugins

Guide to EasySpider for visual no-code web data collection. An agent skill from wentorai/research-plugins.

MITAuto-check passedData & Analytics

Install Easy Spider Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill easy-spider-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins easy-spider-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/scraping/easy-spider-guide .claude/skills/easy-spider-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
easy-spider-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
822 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Guide to EasySpider for visual no-code web data collection. An agent skill from wentorai/research-plugins.

  • Works in 6 steps: Open target page - Enter the URL in… → Select elements - Click on the data… → Define fields - Name each extracted… → …
  • Tasks that involve Web scraping
  • SKILL.md covers Overview, Installation, Core Concepts and Research Use Cases, plus 4 more sections
  • Calls npm and git; reaches github.com

What it does

Easy Spider Guide is an agent skill from wentorai/research-plugins. Guide to EasySpider for visual no-code web data collection

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/easy-spider-guide”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Open target page - Enter the URL in EasySpider's built-in browser
  2. Select elements - Click on the data elements you want to extract
  3. Define fields - Name each extracted element (title, author, date, etc.)
  4. Configure pagination - Click the "next page" button to set up pagination
  5. Set extraction rules - Define how to handle lists, tables, and nested pages
  6. Test and run - Preview results, then execute the full scraping task

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Easy Spider Guide loads about 2.3k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 822 words, ~2,283 tokens.

Download SKILL.mdSave it as .claude/skills/easy-spider-guide/SKILL.md (or your agent's skills folder).
name
easy-spider-guide
description
Guide to EasySpider for visual no-code web data collection

EasySpider Guide

Overview

EasySpider is a visual, no-code web crawler tool with over 44K stars on GitHub. It provides a graphical interface where users design web scraping tasks by interacting directly with target web pages, clicking on elements to extract, and defining navigation flows visually. No programming knowledge is required to build functional scrapers, making it accessible to researchers across all disciplines.

For academic researchers, data collection from web sources is a frequent need but often a technical barrier. Whether gathering publication metadata from journal websites, collecting survey responses from public forums, extracting pricing data for economic research, or archiving web content for digital humanities projects, EasySpider enables researchers to build custom scrapers without writing Python or JavaScript code. The visual approach also makes scrapers easier to maintain and modify when target websites change their structure.

EasySpider runs as a desktop application on Windows, macOS, and Linux. It uses a built-in Chromium browser for rendering, which means it can handle JavaScript-heavy websites, single-page applications, and sites that require user interaction such as clicking buttons, scrolling, or filling forms. Scraped data can be exported as CSV, JSON, or directly to databases.

Installation

Download and Setup
bash
# Download the latest release for your platform from GitHub releases
# https://github.com/NaiboWang/EasySpider/releases

# macOS - download the .dmg file and drag to Applications

# Linux - download the AppImage
chmod +x EasySpider-linux-x86_64.AppImage
./EasySpider-linux-x86_64.AppImage

# Or run from source
git clone https://github.com/NaiboWang/EasySpider.git
cd EasySpider
npm install
npm start
System Requirements
  • Operating system: Windows 10+, macOS 10.15+, or Linux (Ubuntu 18.04+)
  • RAM: 4 GB minimum, 8 GB recommended for complex scraping tasks
  • Disk space: 500 MB for the application plus storage for scraped data
  • Network: stable internet connection for web scraping

Core Concepts

Task Design Workflow

EasySpider follows a visual task design approach with these steps:

  1. Open target page - Enter the URL in EasySpider's built-in browser
  2. Select elements - Click on the data elements you want to extract
  3. Define fields - Name each extracted element (title, author, date, etc.)
  4. Configure pagination - Click the "next page" button to set up pagination
  5. Set extraction rules - Define how to handle lists, tables, and nested pages
  6. Test and run - Preview results, then execute the full scraping task
Element Selection Modes
  • Single element - Click one element to extract that specific item
  • Similar elements - Click two similar items and EasySpider detects the pattern for all matching elements on the page
  • Table mode - Select a table header row to extract entire structured tables
  • Input mode - Define form fields to fill before extracting (useful for search-based data collection)

Research Use Cases

Collecting Publication Metadata

Researchers can use EasySpider to gather publication information from journal websites, conference proceedings pages, or institutional repositories.

Example workflow for scraping a conference proceedings page:

  1. Navigate to the proceedings listing page
  2. Click on the first paper title to mark it as a "title" field
  3. Click on the second paper title; EasySpider recognizes the pattern and selects all titles
  4. Similarly select author names, abstract snippets, and publication dates
  5. If papers span multiple pages, click the "Next" pagination button
  6. Configure "click into each paper" to follow links and extract full abstracts
  7. Run the task and export as CSV
Monitoring Research Funding Opportunities
Task: Daily scan of funding agency websites for new opportunities

Steps configured in EasySpider:
1. Navigate to funding agency announcement page
2. Extract: opportunity title, deadline, funding amount, eligibility
3. Filter: only new announcements (since last check)
4. Schedule: run daily at 8:00 AM
5. Export: append to CSV file, send notification email
Show full SKILL.md (330 more words)Show less
Gathering Economic Data from Public Sources

For economics and social science research, EasySpider can collect publicly available data from government statistics portals, price comparison websites, and public registries.

Task: Collect commodity prices from public market websites

Fields to extract:
- commodity_name: product identifier
- price: current listed price
- unit: measurement unit
- date: listing date
- source_url: page URL for reference

Pagination: navigate through category pages
Schedule: weekly collection
Output: CSV with timestamp for time-series analysis
Digital Humanities Web Archiving
Task: Archive public blog posts for discourse analysis

Configuration:
- Start URL: blog archive page
- Follow: links matching pattern /posts/*
- Extract per page:
  - post_title
  - post_date
  - author_name
  - post_content (full text)
  - comment_count
  - tags/categories
- Pagination: follow archive navigation links
- Output: JSON with full text content

Advanced Features

Conditional Logic

EasySpider supports conditional branches in task flows:

  • If element exists - Check for specific elements before attempting extraction
  • If text contains - Filter items based on content matching
  • Loop control - Set maximum iterations for pagination or nested page visits
Data Cleaning Options

Built-in text processing options can be applied during extraction:

  • Remove HTML tags from extracted text
  • Trim whitespace and normalize spacing
  • Extract numbers from mixed text fields
  • Apply regex patterns to clean specific formats
  • Convert date strings to standardized formats
Handling Dynamic Content

For JavaScript-rendered pages, EasySpider provides options to:

  • Wait for specific elements to appear before extracting
  • Scroll to load lazy-loaded content
  • Click "Load More" buttons automatically
  • Handle infinite scroll pages with configurable scroll limits
Anti-Detection Configuration

For responsible scraping, EasySpider includes options to:

Request configuration:
- Delay between requests: 2-5 seconds (randomized)
- User-Agent rotation: enabled
- Concurrent requests: 1 (sequential for politeness)
- Respect robots.txt: check before scraping
- Rate limiting: max 30 requests per minute

Exporting Research Data

CSV Export

The most common format for researchers. Data is exported with headers matching the field names defined during task design.

csv
title,authors,year,journal,doi,abstract
"Machine Learning in Materials Science","Smith J, Lee K",2025,"Nature Materials","10.1038/xxx","Abstract text here..."
JSON Export

Preserves nested structure for complex extractions:

json
{
  "task_name": "proceedings_scrape",
  "extracted_at": "2026-03-10T14:30:00Z",
  "records": [
    {
      "title": "Machine Learning in Materials Science",
      "authors": ["Smith J", "Lee K"],
      "year": 2025,
      "metadata": {
        "journal": "Nature Materials",
        "doi": "10.1038/xxx"
      }
    }
  ]
}
Database Export

EasySpider can write directly to SQLite databases, which is convenient for subsequent analysis with Python pandas or R.

python
import sqlite3
import pandas as pd

# Read EasySpider output database
conn = sqlite3.connect("easyspider_results.db")
df = pd.read_sql("SELECT * FROM scraped_data", conn)

# Process and analyze
print(f"Total records: {len(df)}")
print(df.describe())
conn.close()

Ethical Web Scraping Guidelines for Researchers

When using EasySpider for research data collection, follow these ethical guidelines:

  • Check robots.txt before scraping any website
  • Respect rate limits and add appropriate delays between requests
  • Review terms of service for target websites
  • Use APIs when available rather than scraping HTML (many services offer research APIs)
  • Minimize data collection to only what is needed for the research question
  • Store data securely especially when collecting personal information
  • Cite data sources in publications and include data collection methodology
  • Obtain IRB approval if scraping involves human subjects data
  • Consider GDPR/privacy regulations when scraping data from EU sources

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/scraping/easy-spider-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Easy Spider Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Easy Spider Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Easy Spider Guide this skillwentorai/research-plugins2981 repos~2.3kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7021 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    702 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 11 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Easy Spider Guide

What does Easy Spider Guide do?

Guide to EasySpider for visual no-code web data collection. An agent skill from wentorai/research-plugins. Easy Spider Guide is an agent skill from wentorai/research-plugins.

When should I use Easy Spider Guide?

Easy Spider Guide fits situations like: tasks that involve Web scraping.

How do I install Easy Spider Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill easy-spider-guide -a claude-code`. Or copy the skill folder (skills/tools/scraping/easy-spider-guide in wentorai/research-plugins) into .claude/skills/easy-spider-guide in your project. Claude Code loads it when a task matches its description.

How do I install Easy Spider Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill easy-spider-guide -a codex`. Or copy the skill folder (skills/tools/scraping/easy-spider-guide in wentorai/research-plugins) into .agents/skills/easy-spider-guide in your project. Codex loads it when a task matches its description.

Can I use Easy Spider Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill easy-spider-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/easy-spider-guide, .gemini/skills/easy-spider-guide, .github/skills/easy-spider-guide and .opencode/skills/easy-spider-guide in your project.

What does Easy Spider Guide need to run?

Going by SKILL.md and its folder, Easy Spider Guide needs the command-line tools its instructions call (npm and git). Our summary lists: Python 3; Node.js.

Does Easy Spider Guide access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org. This is read from the text; nothing was executed.

Is Easy Spider Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Easy Spider Guide use?

Easy Spider Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Easy Spider Guide use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Easy Spider Guide?

Skills that share tags, products or a category with Easy Spider Guide: Tmux (trpc-group/trpc-agent-go, 1.9k stars), Ketch (1broseidon/ketch, 702 stars), Crawl4AI Web Scraping (smallnest/goclaw, 599 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Easy Spider Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.