Agent skill

AI Crawler Access Analysis

by zubair-trabzada in zubair-trabzada/geo-seo-claude

Checks robots.txt, meta tags and HTTP headers to show which AI crawlers can reach a site, then recommends what to allow or block to stay visible in AI search.

MITAuto-check: notesMarketing & SEO

Install AI Crawler Access Analysis

skills CLI
$ npx skills add zubair-trabzada/geo-seo-claude --skill geo-crawlers -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zubair-trabzada/geo-seo-claude geo-crawlers --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zubair-trabzada/geo-seo-claude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/geo-crawlers .claude/skills/geo-crawlers && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
geo-crawlers
GitHub stars
11k
Token cost
~4.7k tokens
SKILL.md length
2,047 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Checks robots.txt, meta tags and HTTP headers to show which AI crawlers can reach a site, then recommends what to allow or block to stay visible in AI search.

  • Works in 7 steps: Fetch and Parse robots.txt → Check Meta Robots Tags → Check HTTP Headers → …
  • Checking whether robots.txt accidentally blocks AI crawlers
  • SKILL.md covers Purpose, Key Insight, Complete AI Crawler Reference and Recommendation Matrix Summary, plus 3 more sections
  • Reaches openai.com and docs.openai.com

What it does

This skill audits whether AI crawlers can read a website at all, treating crawler access as the technical foundation for appearing in AI-generated answers. It inspects robots.txt rules, meta tags and HTTP headers, then produces an access map of which bots are allowed or blocked, with recommendations that balance visibility against control over how content is used.

It carries a crawler reference grouped into tiers by importance for AI search. Each entry lists the operator, user-agent string, purpose, what blocking would cost and a recommendation. The visible tier-one entries are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, with GPTBot noted as possibly feeding model improvement and OAI-SearchBot as search-only. The excerpt is cut off after the first few crawlers, so later tiers are not described here.

When your agent uses it

  • Checking whether robots.txt accidentally blocks AI crawlers
  • Auditing meta tags and headers that restrict AI bots
  • Deciding which AI crawlers to allow for search visibility
  • Preparing a generative engine optimization audit

Example prompts

  • “Check whether example.com blocks GPTBot or OAI-SearchBot in robots.txt.”
  • “Map which AI crawlers can access our site and recommend what to change.”
  • “Our pages never show up in ChatGPT answers. Audit our crawler access first.”

Requirements

  • Network access to fetch the site's robots.txt and headers
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, WebFetch, Write

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Fetch and Parse robots.txt
  2. Check Meta Robots Tags
  3. Check HTTP Headers
  4. Check for AI-Specific Files
  5. Assess JavaScript Rendering Requirements
  6. Parse Content Signals
  7. Detect Cloudflare Managed robots.txt

What it can do on your machine

Read from SKILL.md and the folder at commit 989cae0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • WebFetch
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openai.com
    • docs.openai.com
    • anthropic.com
    • perplexity.ai
    • developer.amazon.com
    • commoncrawl.org
    • contentsignals.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Crawler Access Analysis loads about 4.7k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 2,047 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, WebFetch, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zubair-trabzada/geo-seo-claude at commit 989cae0, republished under its MIT licence (© zubair-trabzada). 2,047 words, ~4,721 tokens.

Download SKILL.mdSave it as .claude/skills/geo-crawlers/SKILL.md (or your agent's skills folder).
name
geo-crawlers
description
AI crawler access analysis. Checks robots.txt, meta tags, and HTTP headers to determine which AI crawlers can access the site. Provides a complete access map and recommendations for maximizing AI visibility while maintaining appropriate control.
allowed-tools
Read, Grep, Glob, Bash, WebFetch, Write

AI Crawler Access Analysis Skill

Purpose

This skill analyzes a website's accessibility to AI crawlers -- the bots that AI companies use to discover, index, and train on web content. If AI crawlers are blocked, the site's content cannot appear in AI-generated responses regardless of its quality. Crawler access is the foundational technical requirement for GEO.

Key Insight

As of early 2026, many websites inadvertently block AI crawlers through overly aggressive robots.txt rules, inherited from legacy SEO configurations. An Originality.ai 2025 study found that over 35% of the top 1,000 websites block at least one major AI crawler, and 5-10% block all AI crawlers. Blocking AI crawlers is the single fastest way to become invisible in AI-generated search results.


Complete AI Crawler Reference

Tier 1: Critical for AI Search Visibility (RECOMMEND: ALLOW)

These crawlers power the AI search products where users actively look for answers. Blocking them directly reduces your visibility in AI-generated responses.

GPTBot
  • Operator: OpenAI
  • User-Agent: GPTBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)
  • Purpose: Fetches content for ChatGPT's web browsing, plugins, and search features. Content accessed by GPTBot may be used to improve OpenAI models.
  • Impact of Blocking: Content will NOT appear in ChatGPT Search results or be accessible when users ask ChatGPT to browse the web. This is the highest-impact AI crawler to allow.
  • Recommendation: ALLOW -- ChatGPT has 300M+ weekly active users as of 2025. Blocking GPTBot removes your content from one of the largest AI search surfaces.
OAI-SearchBot
  • Operator: OpenAI
  • User-Agent: OAI-SearchBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; OAI-SearchBot/1.0; +https://docs.openai.com/bots/overview)
  • Purpose: Specifically powers ChatGPT's search feature. Unlike GPTBot, content accessed by OAI-SearchBot is NOT used for model training -- only for live search results.
  • Impact of Blocking: Content will not appear in ChatGPT's search results even if GPTBot is allowed.
  • Recommendation: ALLOW -- This is a search-only crawler with no training implications. There is no strategic reason to block it.
ChatGPT-User
  • Operator: OpenAI
  • User-Agent: ChatGPT-User
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot)
  • Purpose: Used when a ChatGPT user explicitly asks the model to visit a specific URL. Acts like a browser agent on behalf of the user.
  • Impact of Blocking: ChatGPT cannot visit your pages when users ask it to read or summarize them. This prevents direct user-initiated traffic.
  • Recommendation: ALLOW -- Blocking this bot prevents users who are actively trying to engage with your content from accessing it through ChatGPT.
ClaudeBot
  • Operator: Anthropic
  • User-Agent: ClaudeBot
  • Full User-Agent String: ClaudeBot/1.0; +https://www.anthropic.com/claude-bot
  • Purpose: Fetches web content for Claude's features including web search, citations, and analysis tools.
  • Impact of Blocking: Content will not be accessible to Claude for web search or when users ask Claude to analyze specific URLs.
  • Recommendation: ALLOW -- Claude is a major AI assistant with growing market share. Blocking ClaudeBot reduces your AI search footprint.
PerplexityBot
  • Operator: Perplexity AI
  • User-Agent: PerplexityBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
  • Purpose: Powers Perplexity's AI search engine, which provides sourced answers with direct citations and links back to source pages.
  • Impact of Blocking: Content will not appear in Perplexity search results. Perplexity is one of the best referral traffic sources among AI search products because it always displays source links.
  • Recommendation: ALLOW -- Perplexity drives actual referral traffic and always attributes sources. High-value AI crawler for publishers and businesses.

Tier 2: Important for Broader AI Ecosystem (RECOMMEND: ALLOW)

These crawlers serve large AI platforms or search ecosystems. Allowing them increases your content's reach.

Google-Extended
  • Operator: Google
  • User-Agent: Google-Extended
  • Purpose: Controls whether Google uses your content for Gemini model training and AI Overviews improvement. CRITICAL NOTE: Blocking Google-Extended does NOT affect your Google Search rankings or your appearance in Google Search results. That is controlled by the standard Googlebot.
  • Impact of Blocking: Content may not be used for Gemini training or to improve AI Overviews. However, your content can still appear in AI Overviews based on standard search indexing.
  • Recommendation: ALLOW -- Blocking provides minimal content protection upside while reducing your presence in Google's AI features. Since it does not affect standard search ranking, the only reason to block is philosophical objection to training data usage.
GoogleOther
  • Operator: Google
  • User-Agent: GoogleOther
  • Purpose: Used by Google for various non-search-ranking purposes including research, one-off crawls, and AI-related data collection.
  • Impact of Blocking: Minimal impact on search rankings. May reduce presence in Google's AI research and experimental features.
  • Recommendation: ALLOW -- Low risk, moderate potential benefit for AI feature inclusion.
Applebot-Extended
  • Operator: Apple
  • User-Agent: Applebot-Extended
  • Purpose: Used by Apple to train and improve Apple Intelligence features, Siri, and Apple's AI products. Separate from standard Applebot (which powers Siri search and Spotlight Suggestions).
  • Impact of Blocking: Content may not be used in Apple Intelligence features. Standard Siri and Spotlight functionality is unaffected (controlled by Applebot).
  • Recommendation: ALLOW -- Apple Intelligence is integrated into all Apple devices (2B+ active devices). Presence in Apple's AI features has growing strategic value.
Amazonbot
  • Operator: Amazon
  • User-Agent: Amazonbot
  • Full User-Agent String: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)
  • Purpose: Indexes content for Alexa answers and Amazon's AI features.
  • Impact of Blocking: Content will not appear in Alexa voice responses or Amazon's AI-powered search features.
  • Recommendation: ALLOW -- Relevant for voice search optimization. Lower priority than Tier 1 crawlers but no downside to allowing.
FacebookBot
  • Operator: Meta
  • User-Agent: FacebookBot
  • Purpose: Used by Meta for AI features across Facebook, Instagram, WhatsApp, and Meta AI assistant.
  • Impact of Blocking: Content may not be accessible to Meta AI. Link previews on Facebook/Instagram are handled by a different crawler and are unaffected.
  • Recommendation: ALLOW -- Meta AI is embedded in apps with 3B+ combined users. Growing importance for AI visibility.

Tier 3: Training-Only Crawlers (ALLOW or BLOCK Based on Strategy)

These crawlers are primarily used for AI model training rather than live search features. Blocking them does not affect AI search visibility.

CCBot
  • Operator: Common Crawl (nonprofit)
  • User-Agent: CCBot
  • Full User-Agent String: CCBot/2.0 (https://commoncrawl.org/faq/)
  • Purpose: Builds the Common Crawl dataset, which is used as training data by many AI companies (Google, Meta, Stability AI, and others).
  • Impact of Blocking: Content will not appear in future Common Crawl datasets. Does NOT affect any live AI search product.
  • Recommendation: CONTEXT-DEPENDENT -- Allow if you want maximum long-term AI training presence. Block if you want to control training data usage. No impact on search visibility.
anthropic-ai
  • Operator: Anthropic
  • User-Agent: anthropic-ai
  • Purpose: Used by Anthropic for AI safety research and Claude model training. Separate from ClaudeBot (which powers live features).
  • Impact of Blocking: Content will not be used for Claude training. Does NOT affect Claude's live search or web browsing features (controlled by ClaudeBot).
  • Recommendation: CONTEXT-DEPENDENT -- Similar to CCBot. Allow for training presence, block for training data control. No impact on live AI search.
Bytespider
  • Operator: ByteDance
  • User-Agent: Bytespider
  • Purpose: Used by ByteDance for various AI products including TikTok's AI features and Doubao (their ChatGPT competitor in China).
  • Impact of Blocking: Content will not be used for ByteDance AI products. Minimal impact for Western-market businesses.
  • Recommendation: BLOCK for most Western businesses (aggressive crawling behavior reported, minimal search visibility benefit). ALLOW if targeting Chinese/Asian markets.
cohere-ai
  • Operator: Cohere
  • User-Agent: cohere-ai
  • Purpose: Used by Cohere for model training. Cohere powers enterprise AI solutions and the Coral chat product.
  • Impact of Blocking: Content will not be used for Cohere model training. Minimal direct consumer-facing impact.
  • Recommendation: CONTEXT-DEPENDENT -- Low priority. Allow or block based on general training data stance.

Show full SKILL.md (815 more words)Show less

Recommendation Matrix Summary

CrawlerTierRecommendationReason
GPTBot1ALLOWPowers ChatGPT Search (300M+ users)
OAI-SearchBot1ALLOWSearch-only, no training use
ChatGPT-User1ALLOWUser-initiated browsing
ClaudeBot1ALLOWClaude web search and analysis
PerplexityBot1ALLOWBest referral traffic AI search
Google-Extended2ALLOWGemini features; no search rank impact
GoogleOther2ALLOWGoogle AI research
Applebot-Extended2ALLOWApple Intelligence (2B+ devices)
Amazonbot2ALLOWAlexa and Amazon AI
FacebookBot2ALLOWMeta AI (3B+ app users)
CCBot3ContextTraining data only
anthropic-ai3ContextTraining data only
Bytespider3BLOCKAggressive crawler, low benefit
cohere-ai3ContextTraining data only
Maximum AI Visibility Configuration (robots.txt)

For sites wanting maximum AI search visibility:

# AI Crawlers - ALLOWED for AI search visibility
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: anthropic-ai
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: GoogleOther
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: FacebookBot
Allow: /

# AI Crawlers - BLOCKED (aggressive/low value)
User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

Analysis Procedure

Step 1: Fetch and Parse robots.txt
  1. Use WebFetch to retrieve [domain]/robots.txt.
  2. Parse all User-agent directives and their associated Allow/Disallow rules.
  3. For each AI crawler in the reference list above:
    • Check if there is a specific User-agent block for that crawler
    • Check if there is a wildcard (User-agent: *) block that would apply
    • Determine effective access: Allowed, Blocked, or Not Mentioned (inherits wildcard rules)
  4. Note any Crawl-delay directives that may slow AI crawler access.
  5. Check for Sitemap directives (AI crawlers use these for discovery).
Step 2: Check Meta Robots Tags
  1. For a sample of 5-10 key pages, fetch the HTML and check for:
    • <meta name="robots" content="noindex"> -- blocks all bots
    • <meta name="robots" content="nofollow"> -- prevents link following
    • <meta name="robots" content="noai"> -- emerging tag to block AI use
    • <meta name="robots" content="noimageai"> -- blocks AI image training
    • Bot-specific meta tags: <meta name="GPTBot" content="noindex">
  2. Record any page-level overrides of the robots.txt directives.
Step 3: Check HTTP Headers
  1. For the same sample pages, check response headers for:
    • X-Robots-Tag: noindex -- HTTP header equivalent of meta noindex
    • X-Robots-Tag: noai -- HTTP header to block AI use
    • X-Robots-Tag: noimageai -- blocks AI image training
    • Bot-specific headers: X-Robots-Tag: GPTBot: noindex
  2. Note that HTTP headers override meta tags and apply to non-HTML resources too.
Step 4: Check for AI-Specific Files
  1. Check for /llms.txt (emerging standard for AI crawler guidance).
  2. Check for /.well-known/ai-plugin.json (OpenAI plugin manifest).
  3. Check for /ai.txt (proposed standard, similar to ads.txt for AI).
  4. Record presence/absence and quality of each file.
Step 5: Assess JavaScript Rendering Requirements
  1. Check if the site is a Single Page Application (SPA) or heavily JavaScript-rendered.
  2. AI crawlers vary in their JavaScript rendering capabilities:
    • GPTBot: Limited JS rendering
    • ClaudeBot: Limited JS rendering
    • PerplexityBot: Limited JS rendering
    • Googlebot: Full JS rendering (but Google-Extended inherits this)
  3. If critical content requires JS rendering, flag this as a potential issue.
  4. Check for Server-Side Rendering (SSR) or Static Site Generation (SSG) as mitigations.
Step 6: Parse Content Signals

Using the already-fetched robots.txt from Step 1, scan for Content-Signal: directives (IETF draft draft-romm-aipref-contentsignals).

  1. Scan every line for a line starting with Content-Signal: (case-insensitive).
  2. If found:
    • Parse all key=value pairs (split on , then on =).
    • Validate keys against the known set: ai-train, search, ai-personalization, ai-retrieval.
    • Validate values: only yes and no are valid.
    • Flag any unknown keys or invalid values as a warning — the spec is still an IETF draft.
    • Record the result as Pass and surface parsed values with plain-English meaning.
  3. If absent: record as Recommendation — the site has not declared AI usage preferences.

No additional HTTP request is needed. robots.txt is already fetched in Step 1.

Step 7: Detect Cloudflare Managed robots.txt

Using the already-fetched robots.txt from Step 1, check for a block starting with # BEGIN Cloudflare Managed content (case-insensitive).

  1. If found, Cloudflare is prepending its managed robots.txt to the site's own file. That block disallows GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Amazonbot, Bytespider, CCBot and meta-externalagent.
  2. Record it as a Critical Issue if the site owner's own rules allow AI crawlers. The owner usually doesn't know it's on, because the block never appears in the robots.txt file on their server, only in the live response.
  3. Recommend: in the Cloudflare dashboard, go to Security Settings, filter by "Bot traffic", and turn off "block training in robots.txt" if the site wants to be used by AI. If blocking training is intentional, note it as a deliberate choice, not an issue.

No additional HTTP request is needed. fetch_page.py robots returns this as cloudflare_managed: true.


Output Format

Generate a file called GEO-CRAWLER-ACCESS.md:

markdown
# AI Crawler Access Report: [Domain]

**Analysis Date:** [Date]
**Domain:** [Domain]
**robots.txt Status:** [Found/Not Found/Error]

---

## Crawler Access Summary

| Crawler | Operator | Tier | Status | Impact |
|---|---|---|---|---|
| GPTBot | OpenAI | 1 | [Allowed/Blocked/Not Mentioned] | [Impact description] |
| OAI-SearchBot | OpenAI | 1 | [Status] | [Impact] |
| ChatGPT-User | OpenAI | 1 | [Status] | [Impact] |
| ClaudeBot | Anthropic | 1 | [Status] | [Impact] |
| PerplexityBot | Perplexity | 1 | [Status] | [Impact] |
| Google-Extended | Google | 2 | [Status] | [Impact] |
| GoogleOther | Google | 2 | [Status] | [Impact] |
| Applebot-Extended | Apple | 2 | [Status] | [Impact] |
| Amazonbot | Amazon | 2 | [Status] | [Impact] |
| FacebookBot | Meta | 2 | [Status] | [Impact] |
| CCBot | Common Crawl | 3 | [Status] | [Impact] |
| anthropic-ai | Anthropic | 3 | [Status] | [Impact] |
| Bytespider | ByteDance | 3 | [Status] | [Impact] |
| cohere-ai | Cohere | 3 | [Status] | [Impact] |

## AI Visibility Score: [X]/100

**Tier 1 Access:** [X/5 crawlers allowed]
**Tier 2 Access:** [X/5 crawlers allowed]
**Tier 3 Access:** [X/4 crawlers allowed]

---

## Critical Issues

[List any Tier 1 crawlers that are blocked]

## Recommendations

### Immediate Actions
[Specific robots.txt changes needed]

### robots.txt Recommendation

[Complete recommended robots.txt content for AI crawlers]


### Additional Technical Findings
- **Meta Robots Tags:** [Findings]
- **X-Robots-Tag Headers:** [Findings]
- **JavaScript Rendering:** [Assessment]
- **llms.txt:** [Present/Absent]
- **Sitemap Accessibility:** [Assessment]

### Content Signals (IETF Draft)

**Status:** Present / Absent

<!-- If present: -->
| Signal Key | Value | Meaning |
|---|---|---|
| ai-train | no | Opted out of AI model training |
| search | yes | Permits use in AI-powered search results |

<!-- If absent: -->
**Recommendation:** Add a `Content-Signal:` directive to robots.txt to declare AI usage preferences explicitly. Example:

`Content-Signal: ai-train=no, search=yes, ai-retrieval=yes`

See https://contentsignals.org/ for the full specification.

Scoring for Crawler Access

The AI Crawler Access Score is calculated as:

ComponentWeightScoring
Tier 1 Crawlers Allowed50%20 points per Tier 1 crawler allowed (5 crawlers = 100 points max, scaled to 50)
Tier 2 Crawlers Allowed25%20 points per Tier 2 crawler allowed (5 crawlers = 100 points max, scaled to 25)
No Blanket AI Blocks15%Full points if no User-agent: * Disallow: / and no noai meta tags
AI-Specific Files Present10%5 points for llms.txt, 5 points for sitemap accessible to AI crawlers

Final score = sum of all weighted components, capped at 100.

© zubair-trabzada, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/geo-crawlers of zubair-trabzada/geo-seo-claude.

Open the folder on GitHubat commit 989cae0

Compare with similar skills

AI Crawler Access Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Crawler Access Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Crawler Access Analysis this skillzubair-trabzada/geo-seo-claude11k—~4.7kAutomated safety check: NotesMIT
SEO and GEO Auditdageno-agents/seo-geo-audit176—~2kAutomated safety check: PassMIT
Universal SEO AnalysisAgriciDaniel/claude-seo19k—~4.9kAutomated safety check: PassMIT
SEO AuditAgriciDaniel/codex-seo7992 repos~1.9kAutomated safety check: PassMIT
Full Website SEO AuditAgriciDaniel/claude-seo19k—~2.6kAutomated safety check: PassMIT
Site Launch Checklistsamber/cc-skills227—~13kAutomated safety check: NotesMIT

Similar skills

  • SEO and GEO Audit

    dageno-agents/seo-geo-audit

    Runs one prioritized audit that combines technical SEO, content quality, trust signals, entity clarity and AI search readiness for a page, site or domain.

    176 GitHub stars~2k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • Universal SEO Analysis

    AgriciDaniel/claude-seo

    Hub for site-wide SEO work: audits, technical checks, schema, content quality, local, hreflang and AI-search readiness, run through slash commands.

    19k GitHub stars~4.9k tokensUpdated 6 days ago
    Marketing & SEOAuto-check passed
  • SEO Audit

    AgriciDaniel/codex-seo

    Full website SEO audit with parallel subagent delegation. An agent skill from AgriciDaniel/codex-seo.

    799 GitHub starsUsed in 2 repos~1.9k tokens
    Marketing & SEOAuto-check passed
  • Full Website SEO Audit

    AgriciDaniel/claude-seo

    Runs a full-site SEO audit by crawling the site and delegating to specialist checks, then returns a scored, prioritized report. Meant for site-wide reviews only.

    19k GitHub stars~2.6k tokensUpdated 6 days ago
    Marketing & SEOAuto-check passed
  • Site Launch Checklist

    samber/cc-skills

    Pre-launch checklist for shipping a new website or web app. An agent skill from samber/cc-skills.

    227 GitHub stars~13k tokensUpdated 10 days ago
    Marketing & SEOAuto-check: notes
  • Generative Engine Optimization

    tech-leads-club/agent-skills

    Makes a page or site easier for AI answer engines to find, understand, trust and quote, using metadata, structured data, an llms.txt file and clear page structure.

    7k GitHub stars~2.5k tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed

More from zubair-trabzada/geo-seo-claude

All 16 skills in this repo
  • GEO-First SEO Audit Tool

    zubair-trabzada/geo-seo-claude

    Audits a website for AI search visibility across ChatGPT, Claude, Perplexity and Google AI Overviews while checking traditional SEO, schema and E-E-A-T content quality.

    11k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check: notes
  • GEO Monthly Delta Report

    zubair-trabzada/geo-seo-claude

    Compares a baseline and a current GEO audit for a client, calculates score changes and action item progress, and writes a monthly progress report.

    11k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check: notes
  • AI Citability Scorer

    zubair-trabzada/geo-seo-claude

    Scores how likely AI assistants are to quote passages from a web page and suggests rewrites that make those passages easier to extract.

    11k GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check: notes
  • GEO Service Proposal Generator

    zubair-trabzada/geo-seo-claude

    Builds a client-ready AI-search-optimization proposal from an existing GEO audit, with pricing tiers, an ROI estimate and a markdown document ready to send.

    11k GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • GEO Content E-E-A-T Scorer

    zubair-trabzada/geo-seo-claude

    Scores a page's content against Google's E-E-A-T framework and AI-citability structure, then writes a scored gap-analysis report.

    11k GitHub starsUsed in 2 repos~4k tokens
    Auto-check: notes
  • GEO Prospect Tracker

    zubair-trabzada/geo-seo-claude

    Tracks GEO agency leads and clients through a sales pipeline in a local JSON file, with notes, audit scores, deal values and a pipeline summary.

    11k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about AI Crawler Access Analysis

What does AI Crawler Access Analysis do?

Checks robots.txt, meta tags and HTTP headers to show which AI crawlers can reach a site, then recommends what to allow or block to stay visible in AI search. This skill audits whether AI crawlers can read a website at all, treating crawler access as the technical foundation for appearing in AI-generated answers.txt rules, meta tags and HTTP headers, then produces an access map of which bots are allowed or blocked, with recommendations that balance visibility against control over how content is used.

When should I use AI Crawler Access Analysis?

AI Crawler Access Analysis fits situations like: checking whether robots.txt accidentally blocks AI crawlers; auditing meta tags and headers that restrict AI bots; deciding which AI crawlers to allow for search visibility; preparing a generative engine optimization audit.

How do I install AI Crawler Access Analysis in Claude Code?

Run `npx skills add zubair-trabzada/geo-seo-claude --skill geo-crawlers -a claude-code`. Or copy the skill folder (skills/geo-crawlers in zubair-trabzada/geo-seo-claude) into .claude/skills/geo-crawlers in your project. Claude Code loads it when a task matches its description.

How do I install AI Crawler Access Analysis in Codex?

Run `npx skills add zubair-trabzada/geo-seo-claude --skill geo-crawlers -a codex`. Or copy the skill folder (skills/geo-crawlers in zubair-trabzada/geo-seo-claude) into .agents/skills/geo-crawlers in your project. Codex loads it when a task matches its description.

Can I use AI Crawler Access Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zubair-trabzada/geo-seo-claude --skill geo-crawlers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/geo-crawlers, .gemini/skills/geo-crawlers, .github/skills/geo-crawlers and .opencode/skills/geo-crawlers in your project.

What does AI Crawler Access Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: AI Crawler Access Analysis is instructions for the agent only. Our summary lists: Network access to fetch the site's robots.txt and headers. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, WebFetch, Write.

Does AI Crawler Access Analysis access the network?

SKILL.md names 7 domains. In commands or code: openai.com, docs.openai.com, anthropic.com, perplexity.ai, developer.amazon.com, commoncrawl.org and contentsignals.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is AI Crawler Access Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does AI Crawler Access Analysis use?

AI Crawler Access Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Crawler Access Analysis use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Crawler Access Analysis?

Skills that share tags, products or a category with AI Crawler Access Analysis: SEO and GEO Audit (dageno-agents/seo-geo-audit, 176 stars), Universal SEO Analysis (AgriciDaniel/claude-seo, 19k stars), SEO Audit (AgriciDaniel/codex-seo, 799 stars) and Full Website SEO Audit (AgriciDaniel/claude-seo, 19k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Crawler Access Analysis?

zubair-trabzada (a GitHub user) maintains it in zubair-trabzada/geo-seo-claude, which has 10,982 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 10, 2026.

Source: zubair-trabzada/geo-seo-claude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.