Agent skill

Algo SEO Crawl

by asgard-ai-platform in asgard-ai-platform/skills

Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage.

MITAuto-check passedData & Analytics

Install Algo SEO Crawl

skills CLI
$ npx skills add asgard-ai-platform/skills --skill algo-seo-crawl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install asgard-ai-platform/skills algo-seo-crawl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/asgard-ai-platform/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/algo-seo-crawl .claude/skills/algo-seo-crawl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
algo-seo-crawl
GitHub stars
242
Token cost
~993 tokens
SKILL.md length
400 words
Files
4 (incl. references)
Skills in repo
207
Repo updated
First seen
Licence
MIT

At a glance

Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage.

  • Works in 4 steps: Input Validation → Core Algorithm → Verification → …
  • The user needs to build a site crawler
  • SKILL.md covers Overview, When to Use, Algorithm and Output Format, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Algo SEO Crawl is an agent skill from asgard-ai-platform/skills. Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage. Use this skill when the user needs to build a site crawler, audit website structure, or collect web data systematically — even if they say 'scrape a website', 'crawl all pages', or 'site audit spider'.

Its SKILL.md is about 990 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `examples/sample_scenario.md`, `references/distributed-crawl.md` and `references/url-normalization.md`).

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: 301 open-source coding agent skills across 22 domains — methodology, judgment & gotchas packaged as Claude Agent Skills for the Asgard AI Platform. The licence is MIT.

When your agent uses it

  • The user needs to build a site crawler
  • Audit website structure
  • Collect web data systematically — even if they say scrape a website
  • Crawl all pages

Example prompts

  • “scrape a website”
  • “crawl all pages”
  • “site audit spider”
  • “/algo-seo-crawl”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Input Validation
  2. Core Algorithm
  3. Verification
  4. Output

What it can do on your machine

Read from SKILL.md and the folder at commit 4e7f4f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Algo SEO Crawl loads about 993 tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 400 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~993
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from asgard-ai-platform/skills at commit 4e7f4f8, republished under its MIT licence (© asgard-ai-platform). 400 words, ~993 tokens.

Download SKILL.mdSave it as .claude/skills/algo-seo-crawl/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
algo-seo-crawl
description
Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage. Use this skill when the user needs to build a site crawler, audit website structure, or collect web data systematically — even if they say 'scrape a website', 'crawl all pages', or 'site audit spider'.
metadata.category
WP-35 SEO 演算法
metadata.tags
seo, web-crawler, scraping, site-audit

Web Crawler

Overview

A web crawler systematically traverses web pages by discovering URLs, fetching content, parsing HTML, and storing results. Uses BFS or priority-based frontier management. Performance is I/O-bound, typically limited by politeness constraints rather than compute.

When to Use

Trigger conditions:

  • Building a site audit tool to discover all pages and their link structure
  • Collecting structured data from websites at scale
  • Mapping site architecture for SEO analysis

When NOT to use:

  • When you need data from a single API endpoint (use HTTP client directly)
  • When a sitemap.xml provides all needed URLs (parse sitemap instead)

Algorithm

IRON LAW: Respect robots.txt and Rate Limits
A crawler MUST:
1. Parse and obey robots.txt before crawling any path
2. Enforce crawl-delay (default 1s if unspecified)
3. Identify itself with a descriptive User-Agent
Ignoring these is unethical and will get your IP blocked.
Phase 1: Input Validation

Parse seed URLs, fetch and parse robots.txt for each domain, set crawl scope (same-domain, subdomain, or cross-domain). Gate: Valid seed URLs, robots.txt rules loaded, scope defined.

Phase 2: Core Algorithm
  1. Initialize URL frontier with seed URLs (priority queue or FIFO)
  2. Dequeue URL, check: not visited, allowed by robots.txt, within scope
  3. Fetch page with timeout and retry logic, respect crawl-delay
  4. Parse HTML: extract links (normalize, deduplicate), extract content/metadata
  5. Enqueue discovered URLs, store parsed data
  6. Repeat until frontier empty or limit reached
Phase 3: Verification

Check: no robots.txt violations in crawl log, no duplicate pages stored, all discovered URLs accounted for. Gate: Crawl completed within scope, politeness maintained.

Phase 4: Output

Return site map with pages, link graph, and extracted metadata.

Output Format

json
{
  "pages": [{"url": "...", "status": 200, "title": "...", "links_out": 15, "depth": 2}],
  "metadata": {"pages_crawled": 500, "errors": 12, "duration_seconds": 300, "domain": "example.com"}
}

Examples

Show full SKILL.md (172 more words)Show less
Sample I/O

Input: Seed: "https://example.com", max_depth: 2, max_pages: 100 Expected: Crawl tree with homepage at depth 0, linked pages at depth 1-2, respecting robots.txt

Edge Cases
InputExpectedWhy
robots.txt disallows /Zero pages crawledMust respect full disallow
Redirect loopStop after 5 redirectsPrevent infinite loop
Soft 404 (200 with error page)Flag as soft 404Status code alone is insufficient

Gotchas

  • URL normalization: http://Example.COM/path/ and http://example.com/path are the same URL. Normalize: lowercase host, remove default port, remove trailing slash, sort query params.
  • JavaScript-rendered content: A basic HTTP fetch misses JS-rendered content. Use headless browser (Playwright/Puppeteer) for SPAs.
  • Trap detection: Calendar pages, session IDs in URLs, and infinite pagination create crawler traps. Set max depth and URL pattern limits.
  • Rate limiting yourself: Parallel fetching without per-domain rate limiting will overwhelm small servers. Use per-domain semaphores.
  • Character encoding: Not all pages are UTF-8. Detect encoding from HTTP headers and meta tags; fall back to charset detection libraries.

References

  • For URL normalization rules (RFC 3986), see references/url-normalization.md
  • For distributed crawling architecture, see references/distributed-crawl.md

© asgard-ai-platform, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in algo-seo-crawl of asgard-ai-platform/skills.

  • SKILL.md
  • examples/sample_scenario.md
  • references/distributed-crawl.md
  • references/url-normalization.md

Open the folder on GitHubat commit 4e7f4f8

Compare with similar skills

Algo SEO Crawl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Algo SEO Crawl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Algo SEO Crawl this skillasgard-ai-platform/skills242—~993Automated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7021 repos~3.9kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    702 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from asgard-ai-platform/skills

All 207 skills in this repo
  • Algo Ecom Bm25

    asgard-ai-platform/skills

    Implement BM25 ranking function for e-commerce product search relevance scoring.

    242 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Mfg Cpk

    asgard-ai-platform/skills

    Calculate Cpk process capability index to assess whether a process meets specification requirements.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Price Elasticity

    asgard-ai-platform/skills

    Calculate price elasticity of demand to quantify how price changes affect sales volume.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Bayesian

    asgard-ai-platform/skills

    Apply Bayesian averaging to rank items by combining observed ratings with prior expectations.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Elo

    asgard-ai-platform/skills

    Implement Elo rating system to rank items or players from pairwise comparison outcomes.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Wilson

    asgard-ai-platform/skills

    Calculate Wilson Score confidence intervals for ranking items by positive proportion with sample size correction.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Algo SEO Crawl

What does Algo SEO Crawl do?

Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage. Algo SEO Crawl is an agent skill from asgard-ai-platform/skills. Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage.

When should I use Algo SEO Crawl?

Algo SEO Crawl fits situations like: the user needs to build a site crawler; audit website structure; collect web data systematically — even if they say scrape a website; crawl all pages.

How do I install Algo SEO Crawl in Claude Code?

Run `npx skills add asgard-ai-platform/skills --skill algo-seo-crawl -a claude-code`. Or copy the skill folder (algo-seo-crawl in asgard-ai-platform/skills) into .claude/skills/algo-seo-crawl in your project. Claude Code loads it when a task matches its description.

How do I install Algo SEO Crawl in Codex?

Run `npx skills add asgard-ai-platform/skills --skill algo-seo-crawl -a codex`. Or copy the skill folder (algo-seo-crawl in asgard-ai-platform/skills) into .agents/skills/algo-seo-crawl in your project. Codex loads it when a task matches its description.

Can I use Algo SEO Crawl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgard-ai-platform/skills --skill algo-seo-crawl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/algo-seo-crawl, .gemini/skills/algo-seo-crawl, .github/skills/algo-seo-crawl and .opencode/skills/algo-seo-crawl in your project.

What does Algo SEO Crawl need to run?

SKILL.md names no scripts, command-line tools or credentials: Algo SEO Crawl is instructions for the agent only.

Does Algo SEO Crawl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Algo SEO Crawl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Algo SEO Crawl use?

Algo SEO Crawl is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Algo SEO Crawl use?

About 993 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.

What are the alternatives to Algo SEO Crawl?

Skills that share tags, products or a category with Algo SEO Crawl: Tmux (trpc-group/trpc-agent-go, 1.9k stars), Ketch (1broseidon/ketch, 702 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars) and Crawl4AI Web Scraping (smallnest/goclaw, 599 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Algo SEO Crawl?

asgard-ai-platform (a GitHub organization) maintains it in asgard-ai-platform/skills, which has 242 GitHub stars. The repository holds 207 skills in this directory. The repository was last updated on June 6, 2026.

Source: asgard-ai-platform/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.