Agent skill

Sitemap Analysis and Generation

by AgriciDaniel in AgriciDaniel/claude-seo

Analyzes an existing XML sitemap or generates a new one from industry templates, validating format, URLs, size limits, lastmod accuracy and structure.

MITAuto-check passedMarketing & SEO

Install Sitemap Analysis and Generation

skills CLI
$ npx skills add AgriciDaniel/claude-seo --skill seo-sitemap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AgriciDaniel/claude-seo seo-sitemap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AgriciDaniel/claude-seo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/seo-sitemap .claude/skills/seo-sitemap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
seo-sitemap
GitHub stars
19k
Used in
5 other repos
Token cost
~1.5k tokens
SKILL.md length
620 words
Files
2
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

Analyzes an existing XML sitemap or generates a new one from industry templates, validating format, URLs, size limits, lastmod accuracy and structure.

  • Works in 7 steps: Ask for business type (or auto-detect… → Load industry template from… → Interactive structure planning with user → …
  • Auditing an XML sitemap for broken, redirected or noindexed URLs
  • SKILL.md covers Mode 1: Analyze Existing Sitemap, Mode 2: Generate New Sitemap, Sitemap Format and Error Handling, plus 1 more section
  • Reaches sitemaps.org and google.com

What it does

In analysis mode the skill first discovers candidates with a sitemap_discovery.py helper that reads Sitemap declarations in robots.txt, validates cross-host targets through an SSRF-safe fetch layer and probes common paths when a declared one is stale, so a sitemap is not reported missing too early. Validation covers valid XML, the limit of 50,000 URLs and 50MB uncompressed per file, HTTP 200 for every URL and accurate W3C Datetime lastmod values that reflect real content changes.

Quality signals include a sitemap index above 50,000 URLs, splitting by content type, and no non-canonical, noindexed, redirected or HTTP URLs. Deprecated priority and changefreq tags are flagged as removable, and a severity table maps issues such as oversized files or noindexed URLs to fixes. Image, video and news sitemaps get their own subtype rules, such as up to 1,000 image entries per URL. The excerpt is cut off before the generation mode.

When your agent uses it

  • Auditing an XML sitemap for broken, redirected or noindexed URLs
  • Splitting an oversized sitemap with a sitemap index
  • Generating a new sitemap for a site from a template
  • Checking whether robots.txt points to a working sitemap

Example prompts

  • “Analyze the sitemap for https://example.com and list every URL that does not return 200.”
  • “Generate a sitemap for our blog and split it by content type.”
  • “Check our lastmod values and tell me whether they look accurate.”

Requirements

  • Network access to the site being checked
  • The claude-seo plugin's sitemap_discovery.py helper

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Ask for business type (or auto-detect from existing site)
  2. Load industry template from ../seo-plan/assets/ directory
  3. Interactive structure planning with user
  4. Apply quality gates
  5. Generate valid XML output
  6. Split at whichever comes first: 50,000 URLs or 50MB uncompressed, with sitemap index
  7. Generate STRUCTURE.md documentation

What it can do on your machine

Read from SKILL.md and the folder at commit 4b99de2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are xml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • sitemaps.org
    • google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sitemap Analysis and Generation loads about 1.5k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 620 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AgriciDaniel/claude-seo at commit 4b99de2, republished under its MIT licence (© AgriciDaniel). 620 words, ~1,520 tokens.

Download SKILL.mdSave it as .claude/skills/seo-sitemap/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
seo-sitemap
description
Analyze existing XML sitemaps or generate new ones with industry templates. Validates format, URLs, and structure. Use when user says "sitemap", "generate sitemap", "sitemap issues", or "XML sitemap".
user-invocable
true
argument-hint
[url or generate]
license
MIT
metadata.author
AgriciDaniel
metadata.version
2.4.2
metadata.category
seo

Sitemap Analysis & Generation

Mode 1: Analyze Existing Sitemap

Discover candidates before reporting a sitemap missing:

bash
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json

The helper reads every bounded Sitemap: declaration in robots.txt, validates cross-host targets through the shared SSRF-safe fetch layer, and still probes common paths when a declared sitemap is stale or invalid. Use only entries in found; preserve declared failures as findings instead of treating a robots.txt line alone as proof that a sitemap works.

Validation Checks
  • Valid XML format
  • Per-file limit: ≤50,000 URLs AND ≤50MB uncompressed (whichever is hit first)
  • All URLs return HTTP 200
  • <lastmod> accurate: must be a valid W3C Datetime and reflect the last significant content change (main content, structured data, links, not copyright/boilerplate edits). Google only honours <lastmod> when consistently and verifiably accurate, so warn when values are suspiciously uniform or newer than the page's real content.
  • No deprecated tags: <priority> and <changefreq> are ignored by Google
  • Sitemap referenced in robots.txt
  • Compare crawled pages vs sitemap; flag missing pages
Quality Signals
  • Sitemap index file if >50k URLs
  • Split by content type (pages, posts, images, videos)
  • No non-canonical URLs in sitemap
  • No noindexed URLs in sitemap
  • No redirected URLs in sitemap
  • HTTPS URLs only (no HTTP)
Common Issues
IssueSeverityFix
>50k URLs in single fileCriticalSplit with sitemap index
>50MB uncompressed single fileCriticalSplit with sitemap index
Non-200 URLsHighRemove or fix broken URLs
Noindexed URLs includedHighRemove from sitemap
Redirected URLs includedMediumUpdate to final URLs
All identical lastmodLowUse actual modification dates
Priority/changefreq usedInfoCan remove (ignored by Google)
Extension sitemaps (image / video / news)

Google documents three subtypes with their own rules, validate per-subtype:

  • Image (http://www.google.com/schemas/sitemap-image/1.1): only two valid tags remain, <image:image> and <image:loc> (max 1,000 <image:image> per <url>). <image:caption>/<image:geo_location>/<image:title>/ <image:license> were deprecated (2022), flag as info-level removable.
  • Video: required <video:video> with <video:thumbnail_loc>, <video:title>, <video:description>, plus <video:content_loc> or <video:player_loc>; mRSS also supported. Flag deprecated/removed tags (<video:category>, <video:gallery_loc>, <video:price>, <video:tvshow>, player autoplay/allow_embed) as info-level removable; recheck Google docs before citing a removal date.
  • News: max 1,000 <news:news> per file (not 50,000); include only articles from the last 2 days; required <news:publication>/<news:name>/ <news:language>/<news:publication_date>/<news:title>; submit/discover through Search Console or robots.txt/sitemap index; use Publisher Center only for publication management where relevant. When the news: namespace is detected, override the generic 50k check with the 1,000 cap.
Show full SKILL.md (240 more words)Show less

Mode 2: Generate New Sitemap

Process
  1. Ask for business type (or auto-detect from existing site)
  2. Load industry template from ../seo-plan/assets/ directory
  3. Interactive structure planning with user
  4. Apply quality gates:
    • ⚠️ WARNING at 30+ location pages (require 60%+ unique content)
    • 🛑 HARD STOP at 50+ location pages (require justification)
  5. Generate valid XML output
  6. Split at whichever comes first: 50,000 URLs or 50MB uncompressed, with sitemap index
  7. Generate STRUCTURE.md documentation
Safe Programmatic Pages (OK at scale)

✅ Integration pages (with real setup docs) ✅ Template/tool pages (with downloadable content) ✅ Glossary pages (200+ word definitions) ✅ Product pages (unique specs, reviews) ✅ User profile pages (user-generated content)

Penalty Risk (avoid at scale)

❌ Location pages with only city name swapped ❌ "Best [tool] for [industry]" without industry-specific value ❌ "[Competitor] alternative" without real comparison data ❌ AI-generated pages without human review and unique value

Sitemap Format

Standard Sitemap
xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/page</loc>
    <lastmod>2026-02-07</lastmod>
  </url>
</urlset>
Sitemap Index (for >50k URLs)
xml
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://example.com/sitemap-pages.xml</loc>
    <lastmod>2026-02-07</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemap-posts.xml</loc>
    <lastmod>2026-02-07</lastmod>
  </sitemap>
</sitemapindex>

Error Handling

  • URL unreachable: Report the HTTP status code and suggest checking if the site is live
  • No sitemap found: Run sitemap_discovery.py and report "not found" only when its found list is empty after declared and common candidates are checked
  • Invalid XML format: Report specific parsing errors with line numbers
  • Rate limiting detected: Back off and report partial results with a note about retry timing

Output

For Analysis
  • VALIDATION-REPORT.md: analysis results
  • Issues list with severity
  • Recommendations
For Generation
  • sitemap.xml (or split files with index)
  • STRUCTURE.md: site architecture documentation
  • URL count and organization summary

© AgriciDaniel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/seo-sitemap of AgriciDaniel/claude-seo.

  • SKILL.md
  • LICENSE.txt

Open the folder on GitHubat commit 4b99de2

Used in 7 other repositories

We found 18 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in AgriciDaniel/claude-seo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sitemap Analysis and Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sitemap Analysis and Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sitemap Analysis and Generation this skillAgriciDaniel/claude-seo19k5 repos~1.5kAutomated safety check: PassMIT
GEO-First SEO Audit Toolzubair-trabzada/geo-seo-claude11k—~2.8kAutomated safety check: NotesMIT
SEO Optimizerailabs-393/ai-labs-claude-skills4551 repos~3.2kAutomated safety check: PassMIT
SEO and GEO Auditdageno-agents/seo-geo-audit176—~2kAutomated safety check: PassMIT
SEO Coachakseolabs-seo/seo-coach146—~2.8kAutomated safety check: PassNone
E2E SEO Assistantirinabuht12-oss/marketing-skills4.1k—~1.9kAutomated safety check: PassNone

Similar skills

  • GEO-First SEO Audit Tool

    zubair-trabzada/geo-seo-claude

    Audits a website for AI search visibility across ChatGPT, Claude, Perplexity and Google AI Overviews while checking traditional SEO, schema and E-E-A-T content quality.

    11k GitHub stars~2.8k tokensUpdated today
    Marketing & SEOAuto-check: notes
  • SEO Optimizer

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing HTML/CSS websites for SEO optimization, fixing SEO issues, generating SEO reports, or implementing SEO best practices.

    455 GitHub starsUsed in 1 repo~3.2k tokens
    Marketing & SEOAuto-check passed
  • SEO and GEO Audit

    dageno-agents/seo-geo-audit

    Runs one prioritized audit that combines technical SEO, content quality, trust signals, entity clarity and AI search readiness for a page, site or domain.

    176 GitHub stars~2k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • SEO Coach

    akseolabs-seo/seo-coach

    Beginner-first SEO coaching for people who want to learn by doing one safe, verifiable step at a time.

    146 GitHub stars~2.8k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • E2E SEO Assistant

    irinabuht12-oss/marketing-skills

    Full SEO workflow covering technical audits, content gaps, backlink opportunities, on-page fixes, and content briefs.

    4.1k GitHub stars~1.9k tokensUpdated 17 days ago
    Marketing & SEOAuto-check passed
  • SEO Audit

    andrew-yangy/gru-ai

    Full website SEO audit with parallel subagent delegation. An agent skill from andrew-yangy/gru-ai.

    155 GitHub stars~731 tokensUpdated 7 mo ago
    Marketing & SEOAuto-check passed

More from AgriciDaniel/claude-seo

All 33 skills in this repo
  • Hreflang and International SEO

    AgriciDaniel/claude-seo

    Audits, validates and generates hreflang tags for multi-language and multi-region sites in HTML, HTTP headers or XML sitemaps, flagging common code and return-tag mistakes.

    19k GitHub starsUsed in 5 repos~3.4k tokens
    Auto-check passed
  • Google SEO APIs

    AgriciDaniel/claude-seo

    Pulls real Google data for SEO work: Search Console, PageSpeed Insights, CrUX field data, the Indexing API and GA4 organic traffic, through /seo google commands.

    19k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • SEO Keyword Clustering

    AgriciDaniel/claude-seo

    Clusters keywords by how much their search results overlap and designs a hub-and-spoke content plan with an internal link matrix and an interactive cluster map.

    19k GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • SEO Content Brief Generator

    AgriciDaniel/claude-seo

    Builds research-backed SEO content briefs with competitor scoring, per-section word counts and page-type templates, for new pages or improving existing ones.

    19k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed
  • FLOW SEO Framework

    AgriciDaniel/claude-seo

    Brings the FLOW framework's stage-specific SEO prompts into the agent, from keyword discovery through backlinks, on-page work and conversion to local SEO, loaded on demand.

    19k GitHub starsUsed in 2 repos~1.4k tokens
    Auto-check passed
  • SEO Image Generator

    AgriciDaniel/claude-seo

    Generates Open Graph previews, blog hero images, product photos and infographics for SEO use through Gemini image tools and the banana extension.

    19k GitHub starsUsed in 2 repos~2.1k tokens
    Auto-check passed

Categories

Questions about Sitemap Analysis and Generation

What does Sitemap Analysis and Generation do?

Analyzes an existing XML sitemap or generates a new one from industry templates, validating format, URLs, size limits, lastmod accuracy and structure. txt, validates cross-host targets through an SSRF-safe fetch layer and probes common paths when a declared one is stale, so a sitemap is not reported missing too early. Validation covers valid XML, the limit of 50,000 URLs and 50MB uncompressed per file, HTTP 200 for every URL and accurate W3C Datetime lastmod values that reflect real content changes.

When should I use Sitemap Analysis and Generation?

Sitemap Analysis and Generation fits situations like: auditing an XML sitemap for broken, redirected or noindexed URLs; splitting an oversized sitemap with a sitemap index; generating a new sitemap for a site from a template; checking whether robots.txt points to a working sitemap.

How do I install Sitemap Analysis and Generation in Claude Code?

Run `npx skills add AgriciDaniel/claude-seo --skill seo-sitemap -a claude-code`. Or copy the skill folder (skills/seo-sitemap in AgriciDaniel/claude-seo) into .claude/skills/seo-sitemap in your project. Claude Code loads it when a task matches its description.

How do I install Sitemap Analysis and Generation in Codex?

Run `npx skills add AgriciDaniel/claude-seo --skill seo-sitemap -a codex`. Or copy the skill folder (skills/seo-sitemap in AgriciDaniel/claude-seo) into .agents/skills/seo-sitemap in your project. Codex loads it when a task matches its description.

Can I use Sitemap Analysis and Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgriciDaniel/claude-seo --skill seo-sitemap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/seo-sitemap, .gemini/skills/seo-sitemap, .github/skills/seo-sitemap and .opencode/skills/seo-sitemap in your project.

What does Sitemap Analysis and Generation need to run?

SKILL.md names no scripts, command-line tools or credentials: Sitemap Analysis and Generation is instructions for the agent only. Our summary lists: Network access to the site being checked; The claude-seo plugin's sitemap_discovery.py helper.

Does Sitemap Analysis and Generation access the network?

SKILL.md names 2 domains. In commands or code: sitemaps.org and google.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Sitemap Analysis and Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sitemap Analysis and Generation use?

Sitemap Analysis and Generation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sitemap Analysis and Generation use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sitemap Analysis and Generation?

Skills that share tags, products or a category with Sitemap Analysis and Generation: GEO-First SEO Audit Tool (zubair-trabzada/geo-seo-claude, 11k stars), SEO Optimizer (ailabs-393/ai-labs-claude-skills, 455 stars), SEO and GEO Audit (dageno-agents/seo-geo-audit, 176 stars) and SEO Coach (akseolabs-seo/seo-coach, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sitemap Analysis and Generation?

AgriciDaniel (a GitHub user) maintains it in AgriciDaniel/claude-seo, which has 18,627 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 4, 2026.

Source: AgriciDaniel/claude-seo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.