Agent skill

Site Crawlability

by kostja94 in kostja94/marketing-skills

When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines.

MITAuto-check passedMarketing & SEO

Install Site Crawlability

skills CLI
$ npx skills add kostja94/marketing-skills --skill site-crawlability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kostja94/marketing-skills site-crawlability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kostja94/marketing-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/seo/technical/crawlability .claude/skills/site-crawlability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
site-crawlability
GitHub stars
1k
Token cost
~2k tokens
SKILL.md length
879 words
Files
1
Skills in repo
102
Repo updated
First seen
Licence
MIT

At a glance

When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines.

  • Works in 3 steps: Site structure: Flat vs. deep hierarchy → Framework: Next.js, static, SPA, etc. → Key paths: Sitemap, robots.txt, API,…
  • Wants to improve crawlability
  • SKILL.md covers Scope (Technical SEO), Initial Assessment, Best Practices and Common Issues, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Site Crawlability is an agent skill from kostja94/marketing-skills. When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines. Also use when the user mentions "crawlability," "crawl budget," "orphan pages," "internal links," "site structure," "site crawlability," "infinite scroll," "pagination," "masonry SEO," "AI crawler optimization," "GPTBot crawlability," "ClaudeBot crawlability," or "content not indexed." For internal links, use internal-links.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering Technical SEO. The repository describes itself as: Agent Skills for Marketing — SEO, Social, Influencer & More. 160+ open-source skills for SEO, content, 40+ page types, paid ads, channels, and strategies. Add project context… The licence is MIT.

When your agent uses it

  • Wants to improve crawlability
  • Fix orphan pages
  • Optimize site structure for search engines
  • The user mentions crawlability

Example prompts

  • “crawlability,”
  • “crawl budget,”
  • “orphan pages,”
  • “/site-crawlability”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Site structure: Flat vs. deep hierarchy
  2. Framework: Next.js, static, SPA, etc.
  3. Key paths: Sitemap, robots.txt, API, static assets

What it can do on your machine

Read from SKILL.md and the folder at commit 8dd89c5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.google.com
    • vercel.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Site Crawlability loads about 2k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 879 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kostja94/marketing-skills at commit 8dd89c5, republished under its MIT licence (© kostja94). 879 words, ~1,969 tokens.

Download SKILL.mdSave it as .claude/skills/site-crawlability/SKILL.md (or your agent's skills folder).
name
site-crawlability
description
When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines. Also use when the user mentions "crawlability," "crawl budget," "orphan pages," "internal links," "site structure," "site crawlability," "infinite scroll," "pagination," "masonry SEO," "AI crawler optimization," "GPTBot crawlability," "ClaudeBot crawlability," or "content not indexed." For internal links, use internal-links.
metadata.version
1.2.1

SEO Technical: Crawlability

Guides crawlability improvements: robots, X-Robots-Tag, site structure, and internal linking.

When invoking: On first use, if helpful, open with 1–2 sentences on what this skill covers and why it matters, then provide the main output. On subsequent use or when the user asks to skip, go directly to the main output.

Scope (Technical SEO)

  • Redirect chains & loops: Fix multi-hop redirects; point directly to final URL
  • Broken links (4xx): Fix broken internal/external links; 301 or remove
  • Site architecture: Logical hierarchy; pages within 3–4 clicks from homepage
  • Orphan pages: Add internal links to pages with no incoming links
  • Pagination: Prefer pagination over infinite scroll for crawlability
  • Crawl budget: Reduce waste on duplicates, redirects, low-value URLs (see below)
  • AI crawler optimization: SSR for critical content; URL management; reduce 404/redirect waste (see below)

Initial Assessment

Project context: Read root contextus.md when present and load only the modules relevant to this task. Without Contextus, use available project material or user-provided facts and ask for missing information; do not create a parallel context system.

Identify:

  1. Site structure: Flat vs. deep hierarchy
  2. Framework: Next.js, static, SPA, etc.
  3. Key paths: Sitemap, robots.txt, API, static assets

Best Practices

Redirect Chains & Loops
  • Fix multi-hop redirects; point directly to final URL
  • Loops: URLs redirecting back to themselves; break the cycle
  • Fix broken internal/external links; 301 or remove
  • Audit regularly; update or remove broken links
Site Architecture
PrincipleGuideline
DepthImportant pages within 3–4 clicks from homepage
Orphan pagesAdd internal links to pages with no incoming links; see internal-links for link strategy
HierarchyLogical structure; hub pages link to content
Pagination vs Infinite Scroll

Problem: With infinite scroll, crawlers cannot emulate user behavior (scroll, click "Load more"); content loaded after initial page load is not discoverable. Same applies to masonry + infinite scroll, lazy-loaded lists, and similar patterns.

Solution: Prefer pagination for key content. If keeping infinite scroll, make it search-friendly per Google's recommendations:

RequirementPractice
Component pagesChunk content into paginated pages accessible without JavaScript
Full URLsEach page has unique URL (e.g. ?page=1, ?lastid=567); avoid #1
No overlapEach item listed once in series; no duplication across pages
Direct accessURL works in new tab; no cookie/history dependency
pushState/replaceStateUpdate URL as user scrolls; enables back/forward, shareable links
404 for out-of-bounds?page=999 returns 404 when only 998 pages exist

Reference: Infinite scroll search-friendly recommendations (Google Search Central, 2014)

Pagination (Traditional)
  • Reference links to next/previous pages; rel="prev" / rel="next" where applicable
  • Avoid dynamic-only loading; ensure links in HTML
Crawl Budget

Crawl budget is the number of URLs Googlebot will crawl on your site in a given period. Large sites (10,000+ pages) may waste up to 30% of crawl budget on duplicates, redirects, and low-value URLs.

Waste sourceFix
Duplicate URLsCanonical; consolidate; 301 to preferred
Redirect chainsPoint directly to final URL
Parameter proliferationUse rel="canonical"; consider Clean-param (Yandex)
Low-value pagesnoindex for thin/duplicate; see indexing
Crawl trapsAvoid infinite URL generation (e.g. faceted filters)

Sitemap: Include only indexable, canonical URLs. See xml-sitemap, canonical-tag.

Show full SKILL.md (378 more words)Show less
AI Crawler Optimization

AI crawlers (GPTBot, ClaudeBot, PerplexityBot, etc.) now represent ~28% of Googlebot's crawl volume. Their behavior differs from search engines—optimizing for both improves GEO (AI search visibility). See generative-engine-optimization for GEO strategy. Vercel/MERJ study (Dec 2024):

FactorAI Crawlers (GPTBot, Claude)Googlebot
JavaScriptDo not execute JS; cannot read client-side rendered contentFull JS rendering
404 rate~34% of fetches hit 404s~8%
Redirects~14% of fetches follow redirects~1.5%
Content in initial HTMLJSON, RSC in initial response can be indexedSame

Recommendations for AI crawlability:

PracticeAction
Server-side renderingCritical content in initial HTML. Use SSR, ISR, or SSG. See rendering-strategies for full guide.
URL managementKeep sitemaps updated; use consistent URL patterns; avoid outdated /static/ assets that cause 404s. AI crawlers frequently hit outdated URLs.
RedirectsFix redirect chains; point directly to final URL. AI crawlers waste ~14% of fetches on redirects.
404 handlingFix broken links; remove or redirect outdated URLs. High 404 rates suggest AI crawlers may use stale URL lists.

Reference: The rise of the AI crawler (Vercel, 2024)

Common Issues

IssueCheck
Redirect chainsUpdate links to point directly to final URL
Broken links301 or remove; audit internal and external
Orphan pagesAdd internal links from hub or navigation; see internal-links for strategy
Infinite scrollProvide paginated component pages; or replace with pagination for key content; see above
AI crawlers missing contentEnsure critical content in initial HTML; see rendering-strategies

Output Format

  • Redirect audit: Chains and loops to fix
  • Broken link audit: 4xx links to fix
  • Site structure: Orphan pages, hierarchy
  • Pagination: Implementation for crawlable content
  • AI crawler: SSR/URL/redirect checks if GEO or AI visibility is a goal
  • seo-strategy: SEO workflow; crawlability is Technical phase (P0)
  • website-structure: Plan which pages to build, page priority, structure planning; use before or alongside crawlability audit
  • robots-txt: robots.txt configuration; AI crawler allow/block (GPTBot, ClaudeBot)
  • xml-sitemap: URL discovery; keep updated to reduce AI crawler 404s
  • google-search-console: Index status, Coverage report
  • indexing: Fix indexing issues
  • internal-links: Internal linking best practices
  • masonry: Masonry + infinite scroll has same crawl issue; layout skill references this for SEO
  • generative-engine-optimization: GEO strategy; AI search visibility; crawlability enables AI citation
  • canonical-tag: Canonical reduces crawl budget waste on duplicates
  • rendering-strategies: SSR, SSG, CSR; content in initial HTML; crawler visibility

© kostja94, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/seo/technical/crawlability of kostja94/marketing-skills.

Open the folder on GitHubat commit 8dd89c5

Compare with similar skills

Site Crawlability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Site Crawlability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Site Crawlability this skillkostja94/marketing-skills1k—~2kAutomated safety check: PassMIT
Google SEO APIsAgriciDaniel/claude-seo19k1 repos~4.2kAutomated safety check: PassMIT
SEO Sitemapseranking/seo-skills160—~2.3kAutomated safety check: PassMIT
Hreflang and International SEOAgriciDaniel/claude-seo19k5 repos~3.4kAutomated safety check: PassMIT
GEO-First SEO Audit Toolzubair-trabzada/geo-seo-claude11k—~2.8kAutomated safety check: NotesMIT
SEO Ops Structural Checklisttigerless-labs/seo-ops700—~3.3kAutomated safety check: NotesNone

Similar skills

  • Google SEO APIs

    AgriciDaniel/claude-seo

    Pulls real Google data for SEO work: Search Console, PageSpeed Insights, CrUX field data, the Indexing API and GA4 organic traffic, through /seo google commands.

    19k GitHub starsUsed in 1 repo~4.2k tokens
    Marketing & SEOAuto-check passed
  • SEO Sitemap

    seranking/seo-skills

    Pull a domain's XML sitemap (and sitemap-of-sitemaps), then compare against the most recent SE Ranking website audit.

    160 GitHub stars~2.3k tokensUpdated 3 mo ago
    Marketing & SEOAuto-check passed
  • Hreflang and International SEO

    AgriciDaniel/claude-seo

    Audits, validates and generates hreflang tags for multi-language and multi-region sites in HTML, HTTP headers or XML sitemaps, flagging common code and return-tag mistakes.

    19k GitHub starsUsed in 5 repos~3.4k tokens
    Marketing & SEOAuto-check passed
  • GEO-First SEO Audit Tool

    zubair-trabzada/geo-seo-claude

    Audits a website for AI search visibility across ChatGPT, Claude, Perplexity and Google AI Overviews while checking traditional SEO, schema and E-E-A-T content quality.

    11k GitHub stars~2.8k tokensUpdated today
    Marketing & SEOAuto-check: notes
  • SEO Ops Structural Checklist

    tigerless-labs/seo-ops

    Checks a site against a deterministic set of structural SEO and GEO checks, either running a crawler-eye report against a URL or reviewing code against the same checklist.

    700 GitHub stars~3.3k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check: notes
  • SEO Optimizer

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing HTML/CSS websites for SEO optimization, fixing SEO issues, generating SEO reports, or implementing SEO best practices.

    454 GitHub starsUsed in 1 repo~3.2k tokens
    Marketing & SEOAuto-check passed

More from kostja94/marketing-skills

All 102 skills in this repo
  • Grokipedia Recommendations

    kostja94/marketing-skills

    When the user wants to add recommendations, links, or content to Grokipedia.

    1k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AI Traffic Tracking

    kostja94/marketing-skills

    When the user wants to track AI search traffic in GA4 or GSC.

    1k GitHub stars~959 tokensUpdated 3 days ago
    Auto-check passed
  • Analytics Tracking

    kostja94/marketing-skills

    When the user wants to set up, audit, or optimize analytics tracking (GA4, events, conversions).

    1k GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check passed
  • Brand Visual Generator

    kostja94/marketing-skills

    When the user wants to define, audit, or apply visual identity (typography, colors, spacing, design tokens, frontend aesthetics).

    1k GitHub stars~3k tokensUpdated 3 days ago
    Auto-check passed
  • Canonical Tag

    kostja94/marketing-skills

    When the user wants to configure canonical URLs, fix duplicate content, or consolidate URL signals.

    1k GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Content Strategy

    kostja94/marketing-skills

    When the user wants to plan content for SEO, create content calendar, or build topic clusters.

    1k GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed

Questions about Site Crawlability

What does Site Crawlability do?

When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines. Site Crawlability is an agent skill from kostja94/marketing-skills. When the user wants to improve crawlability, fix orphan pages, or optimize site structure for search engines.

When should I use Site Crawlability?

Site Crawlability fits situations like: wants to improve crawlability; fix orphan pages; optimize site structure for search engines; the user mentions crawlability.

How do I install Site Crawlability in Claude Code?

Run `npx skills add kostja94/marketing-skills --skill site-crawlability -a claude-code`. Or copy the skill folder (skills/seo/technical/crawlability in kostja94/marketing-skills) into .claude/skills/site-crawlability in your project. Claude Code loads it when a task matches its description.

How do I install Site Crawlability in Codex?

Run `npx skills add kostja94/marketing-skills --skill site-crawlability -a codex`. Or copy the skill folder (skills/seo/technical/crawlability in kostja94/marketing-skills) into .agents/skills/site-crawlability in your project. Codex loads it when a task matches its description.

Can I use Site Crawlability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kostja94/marketing-skills --skill site-crawlability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-crawlability, .gemini/skills/site-crawlability, .github/skills/site-crawlability and .opencode/skills/site-crawlability in your project.

What does Site Crawlability need to run?

SKILL.md names no scripts, command-line tools or credentials: Site Crawlability is instructions for the agent only.

Does Site Crawlability access the network?

SKILL.md names 2 domains. As links in the text: developers.google.com and vercel.com. This is read from the text; nothing was executed.

Is Site Crawlability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Site Crawlability use?

Site Crawlability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Site Crawlability use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Site Crawlability?

Skills that share tags, products or a category with Site Crawlability: Google SEO APIs (AgriciDaniel/claude-seo, 19k stars), SEO Sitemap (seranking/seo-skills, 160 stars), Hreflang and International SEO (AgriciDaniel/claude-seo, 19k stars) and GEO-First SEO Audit Tool (zubair-trabzada/geo-seo-claude, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Site Crawlability?

kostja94 (a GitHub user) maintains it in kostja94/marketing-skills, which has 1,025 GitHub stars. The repository holds 102 skills in this directory. The repository was last updated on October 6, 2026.

Source: kostja94/marketing-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.