Agent skill

Crawl4AI SEO Site Crawler

by artwist-polyakov in artwist-polyakov/polyakov-claude-skills

Crawls a site with Crawl4AI to audit titles, meta tags, H1s, canonicals, navigation and internal links, and to compare landing pages and competitor sites.

MITAuto-check passedMarketing & SEO

SKILL.md written in Russian; this summary is our English description.

Install Crawl4AI SEO Site Crawler

skills CLI
$ npx skills add artwist-polyakov/polyakov-claude-skills --skill crawl4ai-seo -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artwist-polyakov/polyakov-claude-skills crawl4ai-seo --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artwist-polyakov/polyakov-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/crawl4ai-seo/skills/crawl4ai-seo .claude/skills/crawl4ai-seo && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawl4ai-seo
GitHub stars
208
Token cost
~1.7k tokens
SKILL.md length
516 words
Files
13 (incl. scripts, references)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Crawls a site with Crawl4AI to audit titles, meta tags, H1s, canonicals, navigation and internal links, and to compare landing pages and competitor sites.

  • Works in 3 steps: Проверь окружение → Определи цель исследования → Собери launch params
  • Running a full SEO inventory of a site's pages
  • SKILL.md covers Что умеет, Config, Workflow and Scripts, plus 6 more sections
  • Runs Python scripts from its folder; calls python3 and uv

What it does

The skill answers what is actually on a site's pages and how the site is built internally. Tasks include a full page inventory (URL, status, title, H1, meta, canonical, word count), an on-page audit for empty or duplicate titles and H1s, broken canonicals and thin content, an internal linking audit that finds orphan pages with no incoming links, a navigation audit of breadcrumbs, menus and hub pages, landing page comparison, and competitor research through the same pipeline. It complements rank and visibility tools rather than replacing them.

Work is organized as jobs. A seed or init script creates a job from launch parameters (target domain, project, label), crawl_batch.py runs the crawl, and report scripts such as build_navigation_report.py and compare_pages.py produce the outputs, with results and the resolved config stored under cache/jobs. Before crawling, doctor.py checks uv, Python, crawl4ai, playwright and the config, and the agent stops rather than faking a crawl if something is missing. The SKILL.md is in Russian and covers Google and Yandex SEO, including Cyrillic URLs.

When your agent uses it

  • Running a full SEO inventory of a site's pages
  • Finding orphan pages and weakly linked sections
  • Auditing breadcrumbs, menus and navigation consistency
  • Comparing a shortlist of landing pages or competitor pages

Example prompts

  • “Crawl example.com and give me an inventory of titles, H1s and canonicals.”
  • “Find orphan pages and weakly linked hubs on our blog.”
  • “Compare our pricing landing page with the three competitor pages I listed, on on-page signals.”

Requirements

  • Python 3 with uv
  • crawl4ai and playwright installed, as checked by doctor.py

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Проверь окружение
  2. Определи цель исследования
  3. Собери launch params

What it can do on your machine

Read from SKILL.md and the folder at commit 8bbeead. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawl4AI SEO Site Crawler loads about 1.7k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 173 tokens; SKILL.md has 516 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from artwist-polyakov/polyakov-claude-skills at commit 8bbeead, republished under its MIT licence (© artwist-polyakov). 516 words, ~1,740 tokens.

Download SKILL.mdSave it as .claude/skills/crawl4ai-seo/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
crawl4ai-seo
description
SEO-краулер сайтов на базе Crawl4AI. Полный аудит страниц: title, meta, H1, canonical, breadcrumbs, навигация, внутренние ссылки. Инвентаризация сайта, навигационный аудит, сравнение лендингов, анализ конкурентов. Работает для Google и Яндекс SEO (Cyrillic URL, коммерческие факторы, региональность). Связка с yandex-search-api, yandex-metrika, yandex-webmaster, scrapedo-web-scraper. Triggers: crawl4ai, seo crawl, site audit, page inventory, site inventory, on-page audit, internal links, internal linking audit, navigation audit, landing comparison, competitor analysis, competitor pages, orphan pages, technical seo, аудит сайта, краулер, перелинковка, навигационный аудит.

crawl4ai-seo

SEO-краулер для аудита сайтов и конкурентного анализа. Отвечает на вопрос: "что реально лежит на страницах и как сайт устроен изнутри".

Не заменяет SERP-инструменты (позиции, видимость) — дополняет их данными о содержимом и структуре страниц.

Что умеет

ЗадачаОписание
Site inventoryПолная инвентаризация страниц: URL, status, title, H1, meta, canonical, word count
On-page auditПоиск проблем: пустые title/H1, дубли заголовков, битые canonical, thin content
Internal linking auditГраф внутренних ссылок, orphan pages (0 входящих), слабо связанные страницы
Navigation auditBreadcrumbs, nav-блоки, menu consistency, hub-страницы без исходящих ссылок
Landing comparisonСравнение shortlist URL по on-page сигналам (title, H1, content, links)
Competitor researchПрогон страниц конкурентов через тот же pipeline, сравнение шаблонов

Config

Глобальные defaults: config/defaults.example.json, локальный override — config/defaults.json. Подробности: config/README.md.

Правило: в defaults не храним конкретный сайт или параметры клиента. Всё site-specific приходит через launch params и фиксируется в cache/jobs/<job_id>/resolved_config.json.

Workflow

Перед любым crawl
  1. Проверь окружение:

    bash
    python3 scripts/doctor.py

    Если crawl4ai/playwright не готовы — остановись на seed/discovery, не имитируй crawl.

  2. Определи цель исследования:

    • Полный аудит сайта → site inventory + navigation audit
    • Аудит лендингов → landing comparison
    • Анализ конкурентов → competitor research
    • Аудит перелинковки → internal linking audit
  3. Собери launch params:

    • target.domain — домен исследования
    • job.project — проект/клиент
    • job.label — метка запуска
Основной pipeline
bash
# 1. Создать job
python3 scripts/init_job.py --domain https://example.com --project client-a --label full-audit

# 2. Собрать seed URL из sitemap/robots
python3 scripts/seed_urls.py --job-id <job_id>

# 3. Запустить crawl
uv run --script scripts/crawl_batch.py --job-id <job_id>

# 4. Построить навигационный отчёт
python3 scripts/build_navigation_report.py \
  --seed-job-id <seed_job_id> \
  --crawl-job-id <crawl_job_id> \
  --report-dir reports/<project>/<label>

# 5. Сравнить страницы
python3 scripts/compare_pages.py --job-id <job_id>
Быстрый вариант (без отдельного init)
bash
python3 scripts/seed_urls.py --domain https://example.com --project client-a --label quick-check
uv run --script scripts/crawl_batch.py --job-id <job_id>

Scripts

ScriptНазначение
doctor.pyПроверка окружения: uv, Python, crawl4ai, playwright, config
init_job.pyСоздание job: job_id, launch_params.json, resolved_config.json
seed_urls.pyDiscovery URL из sitemap/robots/файла → seed.json (работает без crawl4ai)
crawl_batch.pyBatch crawl через crawl4ai → pages.ndjson, links.ndjson, summary.json
compare_pages.pyТабличное сравнение crawled pages по on-page сигналам
build_navigation_report.pyСводный навигационный аудит: orphans, weak hubs, breadcrumbs, menu drift
common.pyShared helpers (не запускать напрямую)

Extracted SEO Signals

На каждую страницу (pages.ndjson) извлекаются:

On-page: title, description, keywords, canonical, robots, og:title, og:description, og:image, H1, headings hierarchy, word count.

Navigation: path depth, breadcrumbs (texts + URLs), nav blocks count, nav links count, nav URLs sample.

Links: internal/external links count. В links.ndjson — полный граф: source → target, anchor text, nofollow, same_domain.

Show full SKILL.md (209 more words)Show less

Navigation Audit Findings

build_navigation_report.py автоматически находит:

ПроблемаЧто значит для SEO
Orphan pages (0 входящих ссылок)Поисковик не найдёт страницу или не передаст ей вес
Weakly linked (1 входящая)Страница получает минимум link equity
Missing breadcrumbsНет навигационной цепочки для бота и пользователя
Breadcrumb inconsistencyРазные trail signatures в одной секции — путаница в иерархии
Weak nav templateHub-страница с подозрительно малым числом nav-ссылок
Menu inconsistencyСтраницы теряют часть common nav targets (шаблон "дрейфует")
Weak hubsКатегория/раздел отдаёт <5 внутренних ссылок
Duplicate titlesНесколько страниц с одинаковым title — каннибализация
Canonical mismatchescanonical указывает не на себя
Technical junklogin-URL, php endpoints, asset-ссылки в навигационном слое
Linked-not-seededСайт линкует URL, которых нет в sitemap

Связка с другими скиллами

СценарийСкилл-партнёрПорядок
SERP → on-page audityandex-search-apiПолучи shortlist из SERP → прогони через crawl4ai
Organic landings → audityandex-metrikaВозьми top organic pages → проверь on-page quality
Индексация + audityandex-webmasterСравни indexed pages с crawled inventory
Anti-block fallbackscrapedo-web-scraperПри блокировке crawl4ai — fallback через Scrape.do

Cache Layout

text
cache/jobs/<job_id>/
  launch_params.json      # raw params этого запуска
  resolved_config.json    # effective config после merge defaults + params
  manifest.json           # operational state и artifact paths
  seed.json               # seed URLs и метаданные discovery
  pages.ndjson            # одна запись на страницу
  links.ndjson            # одна запись на link edge
  summary.json            # компактная сводка
  markdown/               # сохранённый markdown контент

Каждый job изолирован. Параллельные запуски по разным сайтам безопасны.

Output Hygiene

  • stdout: только preview и пути к файлам.
  • Полные данные: только в job files.
  • Для анализа используй rg, head, wc -l по ndjson-файлам, не поднимай browser повторно.

References

© artwist-polyakov, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in plugins/crawl4ai-seo/skills/crawl4ai-seo of artwist-polyakov/polyakov-claude-skills.

  • SKILL.md
  • .gitignore
  • cache/.gitkeep
  • config/README.md
  • config/defaults.example.json
  • references/OUTPUT_SCHEMA.md
  • references/WORKFLOWS.md
  • scripts/build_navigation_report.py
  • scripts/common.py
  • scripts/compare_pages.py
  • scripts/crawl_batch.py
  • scripts/doctor.py
  • scripts/seed_urls.py

Open the folder on GitHubat commit 8bbeead

Compare with similar skills

Crawl4AI SEO Site Crawler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawl4AI SEO Site Crawler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawl4AI SEO Site Crawler this skillartwist-polyakov/polyakov-claude-skills208—~1.7kAutomated safety check: PassMIT
Universal SEO AnalysisAgriciDaniel/claude-seo19k—~4.9kAutomated safety check: PassMIT
SEO Content Auditgooseworks-ai/goose-skills1.2k1 repos~2.8kAutomated safety check: PassMIT
Site ArchitectureAvdLee/RocketSimApp80611 repos~3.3kAutomated safety check: PassCustom licence
SEO Optimizerailabs-393/ai-labs-claude-skills4551 repos~3.2kAutomated safety check: PassMIT
SEO Drift MonitorAgriciDaniel/claude-seo19k1 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Universal SEO Analysis

    AgriciDaniel/claude-seo

    Hub for site-wide SEO work: audits, technical checks, schema, content quality, local, hreflang and AI-search readiness, run through slash commands.

    19k GitHub stars~4.9k tokensUpdated 6 days ago
    Marketing & SEOAuto-check passed
  • SEO Content Audit

    gooseworks-ai/goose-skills

    Comprehensive SEO footprint analysis. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Marketing & SEOAuto-check passed
  • Site Architecture

    AvdLee/RocketSimApp

    When the user wants to plan, map, or restructure their website's page hierarchy, navigation, URL structure, or internal linking.

    806 GitHub starsUsed in 11 repos~3.3k tokens
    Marketing & SEOAuto-check passed
  • SEO Optimizer

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing HTML/CSS websites for SEO optimization, fixing SEO issues, generating SEO reports, or implementing SEO best practices.

    455 GitHub starsUsed in 1 repo~3.2k tokens
    Marketing & SEOAuto-check passed
  • SEO Drift Monitor

    AgriciDaniel/claude-seo

    Captures baselines of a page's SEO-critical elements, compares later snapshots against them and flags regressions by severity, like version control for on-page SEO.

    19k GitHub starsUsed in 1 repo~1.9k tokens
    Marketing & SEOAuto-check passed
  • SEO Audit

    freekmurze/dotfiles

    When the user wants to audit, review, or diagnose SEO issues on their site.

    1k GitHub starsUsed in 34 repos~2.2k tokens
    Marketing & SEOAuto-check passed

More from artwist-polyakov/polyakov-claude-skills

All 21 skills in this repo
  • Agent Deck Sessions

    artwist-polyakov/polyakov-claude-skills

    Launches, monitors and collects results from child AI agent sessions with the agent-deck terminal session manager.

    208 GitHub stars~965 tokensUpdated 3 days ago
    Auto-check passed
  • GitHub Pages Publisher

    artwist-polyakov/polyakov-claude-skills

    Publishes already-built static pages to a GitHub Pages repo under a year, year-month and slug folder layout, optimizes large images and returns the public URL.

    208 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Perplexity Search and Research

    artwist-polyakov/polyakov-claude-skills

    Shell scripts for web search and research through the Perplexity API: raw results, cited answers, background deep research and page fetching, with results cached on disk.

    208 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check: notes
  • Sourcecraft Publisher

    artwist-polyakov/polyakov-claude-skills

    Publish static page artifacts to SourceCraft Sites (Yandex infrastructure, works in Russia), with advisory image optimization and an original-image path.

    208 GitHub stars~990 tokensUpdated 3 days ago
    Auto-check: notes
  • X/Twitter Research via Grok

    artwist-polyakov/polyakov-claude-skills

    Searches X/Twitter through the xAI Grok API to build digests, trend reports and thread analysis, meant to surface post ideas for a Telegram channel.

    208 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check: notes
  • Yandex Metrika

    artwist-polyakov/polyakov-claude-skills

    Яндекс Метрика: детализация трафика по источникам и UTM-меткам, отчёты по конверсиям и поисковым системам, API-сегменты и доступы к счётчикам.

    208 GitHub stars~2.1k tokensUpdated 3 days ago
    Auto-check: notes

Categories

Questions about Crawl4AI SEO Site Crawler

What does Crawl4AI SEO Site Crawler do?

Crawls a site with Crawl4AI to audit titles, meta tags, H1s, canonicals, navigation and internal links, and to compare landing pages and competitor sites. The skill answers what is actually on a site's pages and how the site is built internally. Tasks include a full page inventory (URL, status, title, H1, meta, canonical, word count), an on-page audit for empty or duplicate titles and H1s, broken canonicals and thin content, an internal linking audit that finds orphan pages with no incoming links, a navigation audit of breadcrumbs, menus and hub pages, landing page comparison, and competitor research through the same pipeline.

When should I use Crawl4AI SEO Site Crawler?

Crawl4AI SEO Site Crawler fits situations like: running a full SEO inventory of a site's pages; finding orphan pages and weakly linked sections; auditing breadcrumbs, menus and navigation consistency; comparing a shortlist of landing pages or competitor pages.

How do I install Crawl4AI SEO Site Crawler in Claude Code?

Run `npx skills add artwist-polyakov/polyakov-claude-skills --skill crawl4ai-seo -a claude-code`. Or copy the skill folder (plugins/crawl4ai-seo/skills/crawl4ai-seo in artwist-polyakov/polyakov-claude-skills) into .claude/skills/crawl4ai-seo in your project. Claude Code loads it when a task matches its description.

How do I install Crawl4AI SEO Site Crawler in Codex?

Run `npx skills add artwist-polyakov/polyakov-claude-skills --skill crawl4ai-seo -a codex`. Or copy the skill folder (plugins/crawl4ai-seo/skills/crawl4ai-seo in artwist-polyakov/polyakov-claude-skills) into .agents/skills/crawl4ai-seo in your project. Codex loads it when a task matches its description.

Can I use Crawl4AI SEO Site Crawler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artwist-polyakov/polyakov-claude-skills --skill crawl4ai-seo -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawl4ai-seo, .gemini/skills/crawl4ai-seo, .github/skills/crawl4ai-seo and .opencode/skills/crawl4ai-seo in your project.

What does Crawl4AI SEO Site Crawler need to run?

Going by SKILL.md and its folder, Crawl4AI SEO Site Crawler needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and uv). Our summary lists: Python 3 with uv; crawl4ai and playwright installed, as checked by doctor.py.

Does Crawl4AI SEO Site Crawler access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Crawl4AI SEO Site Crawler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Crawl4AI SEO Site Crawler use?

Crawl4AI SEO Site Crawler is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawl4AI SEO Site Crawler use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Crawl4AI SEO Site Crawler?

Skills that share tags, products or a category with Crawl4AI SEO Site Crawler: Universal SEO Analysis (AgriciDaniel/claude-seo, 19k stars), SEO Content Audit (gooseworks-ai/goose-skills, 1.2k stars), Site Architecture (AvdLee/RocketSimApp, 806 stars) and SEO Optimizer (ailabs-393/ai-labs-claude-skills, 455 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawl4AI SEO Site Crawler?

artwist-polyakov (a GitHub user) maintains it in artwist-polyakov/polyakov-claude-skills, which has 208 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 8, 2026.

Source: artwist-polyakov/polyakov-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.