Agent skill

LLM Crawler Access Check

by davepoon in davepoon/buildwithclaude

Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot.

MITAuto-check passedData & Analytics

Install LLM Crawler Access Check

skills CLI
$ npx skills add davepoon/buildwithclaude --skill llm-crawler-access-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davepoon/buildwithclaude llm-crawler-access-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/all-skills/skills/llm-crawler-access-check .claude/skills/llm-crawler-access-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-crawler-access-check
GitHub stars
3.6k
Token cost
~1.5k tokens
SKILL.md length
799 words
Files
1
Skills in repo
246
Repo updated
First seen
Licence
MIT

At a glance

Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot.

  • Works in 3 steps: Fetch → Resolve each agent → Report
  • Someone asks whether AI bots are blocked
  • SKILL.md covers Scope, Procedure, Three mistakes this check… and If the user asks whether they…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

LLM Crawler Access Check is an agent skill from davepoon/buildwithclaude. Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. Use when someone asks whether AI bots are blocked, whether to allow or block GPTBot, why a site never appears in AI answers, or wants a robots.txt review for AI crawlers. Reads only robots.txt, then returns a per-agent allow/block table, the exact rule responsible for each verdict, and the precise lines to change. Distinguishes training crawlers from the search…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping, Technical SEO and Web search. It works with OpenAI and Perplexity. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.

When your agent uses it

  • Someone asks whether AI bots are blocked
  • Whether to allow
  • Why a site never appears in AI answers
  • Wants a robots.txt review for AI crawlers

Example prompts

  • “/llm-crawler-access-check”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Fetch
  2. Resolve each agent
  3. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 616deb5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • maxaeo.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Crawler Access Check loads about 1.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 799 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davepoon/buildwithclaude at commit 616deb5, republished under its MIT licence (© davepoon). 799 words, ~1,523 tokens.

Download SKILL.mdSave it as .claude/skills/llm-crawler-access-check/SKILL.md (or your agent's skills folder).
name
llm-crawler-access-check
description
Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. Use when someone asks whether AI bots are blocked, whether to allow or block GPTBot, why a site never appears in AI answers, or wants a robots.txt review for AI crawlers. Reads only robots.txt, then returns a per-agent allow/block table, the exact rule responsible for each verdict, and the precise lines to change. Distinguishes training crawlers from the search crawlers that actually control citations.
category
security
license
MIT

AI crawler access check

One wrong line in robots.txt removes a site from AI answers completely, and no amount of content work can compensate. This check takes under a minute and should run before any other AI-visibility work.

Scope

Read https://<domain>/robots.txt and nothing else. Do not crawl the site, do not attempt to access disallowed paths, and do not bypass any access control. This is a read of one public file.

Procedure

1. Fetch

Fetch https://<domain>/robots.txt.

  • 404 or empty - everything is allowed by default. Say so; that is a valid and often correct configuration. Stop and report.
  • Non-200 other than 404, or unreachable - report the status code and stop. Do not guess at contents.
  • Served as HTML (a soft 404 returning the site's error page) - flag it. Crawlers may parse it as garbage. This is itself a finding.
2. Resolve each agent

For each agent below, apply standard robots.txt matching: the most specific User-agent group that names the agent wins, and * applies only when no group names it. Within the winning group, the longest matching path rule wins, and Allow beats Disallow on an equal-length match.

AgentOperatorPurposeWhat blocking it actually costs
OAI-SearchBotOpenAIsearch indexcitations in ChatGPT Search
ChatGPT-UserOpenAIlive fetch during a chatthe model cannot open your page when a user asks about it
GPTBotOpenAItrainingbackground model knowledge, not search citations
PerplexityBotPerplexitysearch indexPerplexity citations
Perplexity-UserPerplexitylive fetch during a querylive page reads
ClaudeBotAnthropicindex and trainingAnthropic-side retrieval
GooglebotGooglemain indexAI Overviews and AI Mode, plus normal search
Google-ExtendedGoogleGemini grounding and trainingGemini grounding only - not AI Overviews
BingbotMicrosoftBing indexMicrosoft Copilot, which rides the Bing index
ApplebotAppleindexApple search surfaces
Applebot-ExtendedAppletrainingApple Intelligence training only
CCBotCommon Crawlopen crawl corpusan input to many downstream models

Crawler names change and new ones appear. Before finalizing, check each operator's own published crawler documentation for agents added or renamed since this list was written, and include them. State which list you used.

3. Report

Produce a table with one row per agent and exactly these columns:

Agent | Verdict (ALLOWED / BLOCKED / PARTIAL) | Rule responsible | Impact

  • Rule responsible must quote the literal line from robots.txt, or say no matching rule - allowed by default. Never state a verdict without the line that produced it.
  • PARTIAL means important paths are disallowed while the site root is allowed. Name the disallowed paths.

Then give:

  • Verdict - one sentence: is this site reachable by AI answer engines, or not?
  • What to change - the exact robots.txt lines to add, remove, or edit, as a code block the user can paste. If nothing needs to change, say that plainly rather than inventing work.
  • What this check did not cover - robots.txt is only the first gate. Server-side blocking by WAF, CDN bot rules, IP reputation, or Cloudflare bot management can block a crawler that robots.txt allows, and none of that is visible in this file. Say so every time.
Show full SKILL.md (295 more words)Show less

Three mistakes this check exists to catch

  1. Blocking GPTBot to opt out of training, and assuming that is the whole story. It is not. OAI-SearchBot governs whether a site can be cited in ChatGPT Search, and it is a separate agent with a separate rule. Blocking one does not block the other, in either direction.
  2. Blocking Google-Extended to stay out of AI Overviews. It does not do that. AI Overviews and AI Mode are built on the normal Googlebot index. Blocking Google-Extended opts out of Gemini grounding and training and has no effect on AI Overviews. To leave AI Overviews, the mechanism is the nosnippet, max-snippet, or data-nosnippet family, and it costs normal search snippets too. Say that tradeoff out loud rather than letting the user discover it later.
  3. A blanket User-agent: * / Disallow: / inherited from a staging config, a bot-mitigation template, or a security hardening guide. This is common and almost always unintentional on a production marketing site.

If the user asks whether they should block AI crawlers

Do not answer with a recommendation. Lay out the tradeoff and let them decide: allowing search crawlers is what makes citation possible, allowing training crawlers affects model knowledge but not citation, and the two decisions are independent. Publishers with a licensing position and companies that want to be recommended by AI assistants land in different places, and both are legitimate.


About

Maintained by MaxAEO — https://maxaeo.ai — which works on AI answer-engine visibility. The crawler matrix used here is kept current against each operator's own published crawler documentation; where an agent has no official documentation, this skill says so rather than guessing.

This check is free, read-only, and runs on one public file. It does not require an account, an API key, or any paid service.

© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/all-skills/skills/llm-crawler-access-check of davepoon/buildwithclaude.

Open the folder on GitHubat commit 616deb5

Compare with similar skills

LLM Crawler Access Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Crawler Access Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Crawler Access Check this skilldavepoon/buildwithclaude3.6k—~1.5kAutomated safety check: PassMIT
Agent Readiness Auditindranilbanerjee/digital-marketing-pro8591 repos~3.9kAutomated safety check: PassMIT
Money SEOiamzifei/show-me-the-money1k—~4.4kAutomated safety check: PassCustom licence
SEO AgiLeoYeAI/openclaw-master-skills2.2k—~6.4kAutomated safety check: NotesMIT
Nuxt Geo Best Practicesvinayakkulkarni/nxui212—~1.9kAutomated safety check: PassMIT
Brightdata SDK JSbrightdata/skills264—~3kAutomated safety check: PassMIT

Similar skills

  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended)…

    859 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Money SEO

    iamzifei/show-me-the-money

    SEO and GEO (Generative Engine Optimization) for organic traffic and AI search visibility.

    1k GitHub stars~4.4k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • SEO Agi

    LeoYeAI/openclaw-master-skills

    Write SEO pages that rank in Google AND get cited by LLMs (ChatGPT, Perplexity, Claude).

    2.2k GitHub stars~6.4k tokensUpdated 2 mo ago
    Marketing & SEOAuto-check: notes
  • Nuxt Geo Best Practices

    vinayakkulkarni/nxui

    Nuxt GEO (Generative Engine Optimization) guidelines for getting cited by ChatGPT, Perplexity, Claude, Google AI Overviews, and Gemini.

    212 GitHub stars~1.9k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • SEO Profound

    AgriciDaniel/claude-seo

    Profound LLM citation tracker (extension). An agent skill from AgriciDaniel/claude-seo.

    19k GitHub starsUsed in 1 repo~441 tokens
    Marketing & SEOAuto-check passed

More from davepoon/buildwithclaude

All 246 skills in this repo
  • iOS Hig Design Guide

    davepoon/buildwithclaude

    Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.

    3.6k GitHub stars~735 tokensUpdated today
    Auto-check passed
  • Video Downloader

    davepoon/buildwithclaude

    Download YouTube videos with customizable quality and format options.

    3.6k GitHub starsUsed in 1 repo~871 tokens
    Auto-check passed
  • Qwen Vision

    davepoon/buildwithclaude

    A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

    3.6k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Atlas Cloud Media

    davepoon/buildwithclaude

    Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.

    3.6k GitHub stars~852 tokensUpdated today
    Auto-check passed
  • Browser Extension Launch

    davepoon/buildwithclaude

    面向没有编程经验的用户,把想法做成可试用的浏览器插件,并完成检查、商店材料、审核提交和上线验证;也用于继续已有插件、排错和发布新版。用户说“帮我做个插件”“把插件上架”“继续我的插件”时使用。普通网站开发、仅查询插件知识不触发。

    3.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Slack Gif Creator

    davepoon/buildwithclaude

    Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.

    3.6k GitHub starsUsed in 12 repos~4.3k tokens
    Auto-check passed

Questions about LLM Crawler Access Check

What does LLM Crawler Access Check do?

Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. LLM Crawler Access Check is an agent skill from davepoon/buildwithclaude.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot.

When should I use LLM Crawler Access Check?

LLM Crawler Access Check fits situations like: someone asks whether AI bots are blocked; whether to allow; why a site never appears in AI answers; wants a robots.txt review for AI crawlers.

How do I install LLM Crawler Access Check in Claude Code?

Run `npx skills add davepoon/buildwithclaude --skill llm-crawler-access-check -a claude-code`. Or copy the skill folder (plugins/all-skills/skills/llm-crawler-access-check in davepoon/buildwithclaude) into .claude/skills/llm-crawler-access-check in your project. Claude Code loads it when a task matches its description.

How do I install LLM Crawler Access Check in Codex?

Run `npx skills add davepoon/buildwithclaude --skill llm-crawler-access-check -a codex`. Or copy the skill folder (plugins/all-skills/skills/llm-crawler-access-check in davepoon/buildwithclaude) into .agents/skills/llm-crawler-access-check in your project. Codex loads it when a task matches its description.

Can I use LLM Crawler Access Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill llm-crawler-access-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-crawler-access-check, .gemini/skills/llm-crawler-access-check, .github/skills/llm-crawler-access-check and .opencode/skills/llm-crawler-access-check in your project.

What does LLM Crawler Access Check need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Crawler Access Check is instructions for the agent only.

Does LLM Crawler Access Check access the network?

SKILL.md names 1 domain. As links in the text: maxaeo.ai. This is read from the text; nothing was executed.

Is LLM Crawler Access Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Crawler Access Check use?

LLM Crawler Access Check is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Crawler Access Check use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Crawler Access Check?

Skills that share tags, products or a category with LLM Crawler Access Check: Agent Readiness Audit (indranilbanerjee/digital-marketing-pro, 859 stars), Money SEO (iamzifei/show-me-the-money, 1k stars), SEO Agi (LeoYeAI/openclaw-master-skills, 2.2k stars) and Nuxt Geo Best Practices (vinayakkulkarni/nxui, 212 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Crawler Access Check?

davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,605 GitHub stars. The repository holds 246 skills in this directory. The repository was last updated on October 9, 2026.

Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.