Agent skill

Site Content Catalog

by gooseworks-ai in gooseworks-ai/goose-skills

Crawl a website's sitemap and blog index to build a complete content inventory.

MITAuto-check passedWriting & Content

Install Site Content Catalog

skills CLI
$ npx skills add gooseworks-ai/goose-skills --skill site-content-catalog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gooseworks-ai/goose-skills site-content-catalog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gooseworks-ai/goose-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/seo/capabilities/site-content-catalog .claude/skills/site-content-catalog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
site-content-catalog
GitHub stars
1.2k
Used in
1 other repo
Token cost
~1.9k tokens
SKILL.md length
618 words
Files
3 (incl. scripts)
Skills in repo
273
Repo updated
First seen
Licence
MIT

At a glance

Crawl a website's sitemap and blog index to build a complete content inventory.

  • Works in 5 steps: Discover All Pages → Classify Each Page → Analyze Publishing Patterns → …
  • Tasks that involve Content strategy
  • SKILL.md covers Quick Start, Inputs, Cost and Process, plus 2 more sections
  • Runs Python scripts from its folder; calls python3 and pip; needs APIFY_API_TOKEN

What it does

Site Content Catalog is an agent skill from gooseworks-ai/goose-skills. Crawl a website's sitemap and blog index to build a complete content inventory. Lists every page with URL, title, publish date, content type, and topic cluster. Groups content by category and topic. Optionally deep-reads top N pages for quality analysis and funnel stage tagging. Use before SEO audits, content gap analysis, or brand voice extraction.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/catalog_content.py` and `skill.meta.json`).

It sits in Writing & Content, covering Content strategy, Web scraping and Brand voice and tone. The repository describes itself as: Library of Growth & GTM skills + data APIs for Claude Code, Codex, Cursor to run ads, social, content, lead gen, seo and data scraping. The licence is MIT.

When your agent uses it

  • Tasks that involve Content strategy
  • Tasks that involve Web scraping
  • Tasks that involve Brand voice and tone

Example prompts

  • “/site-content-catalog”

Requirements

  • Python 3
  • A credential in APIFY_API_TOKEN

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Discover All Pages
  2. Classify Each Page
  3. Analyze Publishing Patterns
  4. Deep Analysis (Optional)
  5. Output

What it can do on your machine

Read from SKILL.md and the folder at commit c650c6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • APIFY_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Site Content Catalog loads about 1.9k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 618 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gooseworks-ai/goose-skills at commit c650c6d, republished under its MIT licence (© gooseworks-ai). 618 words, ~1,882 tokens.

Download SKILL.mdSave it as .claude/skills/site-content-catalog/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
site-content-catalog
description
Crawl a website's sitemap and blog index to build a complete content inventory. Lists every page with URL, title, publish date, content type, and topic cluster. Groups content by category and topic. Optionally deep-reads top N pages for quality analysis and funnel stage tagging. Use before SEO audits, content gap analysis, or brand voice extraction.
tags
content, seo

Site Content Catalog

Crawl a website's sitemap and blog to build a complete content inventory — every page cataloged with URL, title, date, content type, and topic cluster. Groups content by category, identifies publishing patterns, and optionally deep-analyzes top pages.

Quick Start

bash
# Basic content inventory
python3 scripts/catalog_content.py --domain "example.com"

# With deep analysis of top 20 pages
python3 scripts/catalog_content.py --domain "example.com" --deep-analyze 20

# Output to specific file
python3 scripts/catalog_content.py --domain "example.com" --output content-inventory.json

Inputs

ParameterRequiredDefaultDescription
domainYes—Domain to catalog (e.g., "example.com")
deep-analyzeNo0Number of top pages to deep-read for content analysis
outputNostdoutPath to save JSON output
include-non-blogNotrueAlso catalog landing pages, docs, etc. (not just blog)

Cost

  • Sitemap/RSS crawling: Free (direct HTTP requests)
  • Apify sitemap extractor (fallback): ~$0.50 per site
  • Deep analysis: Free (WebFetch on individual pages)

Process

Phase 1: Discover All Pages

The script attempts multiple methods to find all pages on a site, in order:

A) Sitemap.xml
  1. Fetch https://[domain]/sitemap.xml
  2. If it's a sitemap index, recursively fetch all child sitemaps
  3. Common alternate locations: /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml
  4. Check robots.txt for Sitemap: directives
B) RSS/Atom Feeds
  1. Check /feed, /rss, /atom.xml, /blog/feed, etc.
  2. Extract posts with titles, dates, and URLs
  3. RSS typically only surfaces recent content (last 10-50 posts)
C) Blog Index Crawl
  1. Fetch /blog, /resources, /insights, /news, /articles
  2. Extract links from the page
  3. Follow pagination if present (/blog/page/2, ?page=2, etc.)
D) Site: Search (fallback)
  1. WebSearch: site:[domain] to estimate total indexed pages
  2. WebSearch: site:[domain]/blog to find blog content
  3. WebSearch: site:[domain] intitle: to discover page title patterns
E) Apify Sitemap Extractor (fallback for JS-heavy sites)
  • Actor: onescales/sitemap-url-extractor
  • Use when sitemap.xml is missing and the site is JS-rendered
Phase 2: Classify Each Page

For each discovered URL, classify by:

Content Type

Classify based on URL patterns and page titles:

TypeURL PatternsExamples
blog-post/blog/, /posts/, /articles/How-to guides, opinion pieces
case-study/case-study/, /customers/, /success-stories/Customer stories
comparison/vs/, /compare/, /alternative/X vs Y pages
landing-page/solutions/, /use-cases/, /for-/Product marketing pages
docs/docs/, /help/, /documentation/, /api/Technical documentation
changelog/changelog/, /releases/, /whats-new/Product updates
pricing/pricing/Pricing page
about/about/, /team/, /careers/Company pages
legal/privacy/, /terms/, /security/Legal/compliance
resource/resources/, /guides/, /ebooks/, /webinars/Gated/downloadable content
glossary/glossary/, /dictionary/, /terms/SEO glossary pages
integration/integrations/, /apps/, /marketplace/Integration pages
other—Anything else
Show full SKILL.md (257 more words)Show less
Topic Cluster

Group by extracting topic signals from URL slugs and titles:

  • Extract keywords from URL path segments
  • Group similar keywords into clusters (e.g., "aws-cost", "cloud-spending", "finops" → "Cloud Cost Management")
  • Use simple keyword co-occurrence for clustering
Phase 3: Analyze Publishing Patterns

From the dated content (primarily blog posts):

  • Total content pieces by type
  • Publishing frequency: Posts per month over last 12 months
  • Trend: Increasing, decreasing, or stable output
  • Recency: Date of most recent publish
  • Author diversity: Unique authors (if extractable from RSS)
Phase 4: Deep Analysis (Optional)

If --deep-analyze N is specified, fetch the top N pages (prioritizing blog posts) and extract:

  • Word count (approximate)
  • Target keyword (inferred from title + H1 + URL)
  • Funnel stage: TOFU (awareness), MOFU (consideration), BOFU (decision)
  • Content depth: Shallow (<500 words), Medium (500-1500), Deep (1500+)
  • Has images/video: Boolean
  • Has CTA: Boolean (detected by common CTA patterns)
  • Internal links count
Phase 5: Output
JSON Output (default)
json
{
  "domain": "example.com",
  "crawl_date": "2026-02-25",
  "total_pages": 347,
  "discovery_methods": ["sitemap.xml", "rss"],
  "pages": [
    {
      "url": "https://example.com/blog/reduce-aws-costs",
      "title": "How to Reduce Your AWS Bill by 40%",
      "date": "2025-11-15",
      "type": "blog-post",
      "topic_cluster": "Cloud Cost Optimization",
      "deep_analysis": {
        "word_count": 2100,
        "target_keyword": "reduce aws costs",
        "funnel_stage": "TOFU",
        "content_depth": "deep",
        "has_images": true,
        "has_cta": true
      }
    }
  ],
  "summary": {
    "by_type": {"blog-post": 89, "landing-page": 23, "case-study": 12, ...},
    "by_topic": {"Cloud Cost Optimization": 34, "FinOps": 18, ...},
    "publishing_cadence": {
      "posts_per_month_avg": 4.2,
      "trend": "increasing",
      "most_recent": "2026-02-20"
    }
  }
}
Markdown Summary (also generated)
markdown
# Content Inventory: example.com
**Crawled:** 2026-02-25 | **Total pages:** 347

## Content by Type
| Type | Count | % |
|------|-------|---|
| Blog Posts | 89 | 25.6% |
| Landing Pages | 23 | 6.6% |
| ...

## Content by Topic Cluster
| Topic | Posts | Most Recent |
|-------|-------|-------------|
| Cloud Cost Optimization | 34 | 2026-02-20 |
| ...

## Publishing Cadence
- Average: 4.2 posts/month
- Trend: Increasing (3.1 → 5.4 over last 6 months)
- Most recent: 2026-02-20

## Full Catalog
| # | Date | Type | Topic | Title | URL |
|---|------|------|-------|-------|-----|
| 1 | 2026-02-20 | blog-post | Cloud Cost | How to Reduce... | https://... |

Tips

  • Sitemap.xml is the best source. Most well-maintained sites have one. If missing, it's itself an SEO signal (negative).
  • RSS only shows recent content. If you need the full catalog, sitemap is essential. RSS is supplementary.
  • Deep analysis is optional but valuable. Use it when feeding into brand-voice-extractor or when you need funnel stage mapping.
  • JS-rendered sites may need the Apify fallback. Signs: sitemap.xml returns HTML, or blog page returns mostly JavaScript.
  • Combine with seo-domain-analyzer to overlay traffic data on the content inventory — see which content actually performs.

Dependencies

  • Python 3.8+
  • requests library (pip install requests)
  • APIFY_API_TOKEN env var (only for Apify fallback mode)

© gooseworks-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/seo/capabilities/site-content-catalog of gooseworks-ai/goose-skills.

  • SKILL.md
  • scripts/catalog_content.py
  • skill.meta.json

Open the folder on GitHubat commit c650c6d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in gooseworks-ai/goose-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Site Content Catalog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Site Content Catalog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Site Content Catalog this skillgooseworks-ai/goose-skills1.2k1 repos~1.9kAutomated safety check: PassMIT
Caption Writerstevenflanagan1/social-ai-team244—~2.9kAutomated safety check: PassNone
SEO Coachakseolabs-seo/seo-coach146—~2.8kAutomated safety check: PassNone
E2E SEO Assistantirinabuht12-oss/marketing-skills4.1k—~1.9kAutomated safety check: PassNone
Brand Voice Content Creatordavila7/claude-code-templates33k3 repos~1.9kAutomated safety check: PassMIT
Crawl4AI SEO Site Crawlerartwist-polyakov/polyakov-claude-skills208—~1.7kAutomated safety check: PassMIT

Similar skills

  • Caption Writer

    stevenflanagan1/social-ai-team

    Writes on-brand social media captions for SMBs. An agent skill from stevenflanagan1/social-ai-team.

    244 GitHub stars~2.9k tokensUpdated 4 days ago
    Writing & ContentAuto-check passed
  • SEO Coach

    akseolabs-seo/seo-coach

    Beginner-first SEO coaching for people who want to learn by doing one safe, verifiable step at a time.

    146 GitHub stars~2.8k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • E2E SEO Assistant

    irinabuht12-oss/marketing-skills

    Full SEO workflow covering technical audits, content gaps, backlink opportunities, on-page fixes, and content briefs.

    4.1k GitHub stars~1.9k tokensUpdated 17 days ago
    Marketing & SEOAuto-check passed
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    33k GitHub starsUsed in 3 repos~1.9k tokens
    Writing & ContentAuto-check passed
  • Crawl4AI SEO Site Crawler

    artwist-polyakov/polyakov-claude-skills

    Crawls a site with Crawl4AI to audit titles, meta tags, H1s, canonicals, navigation and internal links, and to compare landing pages and competitor sites.

    208 GitHub stars~1.7k tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed
  • SEO Site Audit

    spronta/crawlie

    Run a complete technical SEO + AI-search audit of a website with crawlie.

    114 GitHub stars~966 tokensUpdated 2 mo ago
    Marketing & SEOAuto-check passed

More from gooseworks-ai/goose-skills

All 273 skills in this repo
  • Reddit Post Finder

    gooseworks-ai/goose-skills

    Scrape and search Reddit posts using Apify. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Image Fal

    gooseworks-ai/goose-skills

    Generate or edit an image via any FAL image model (nano-banana edit, gpt-image, flux, ...), ROUTED THROUGH THE fal-proxy so it bills the Ads agent.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Render Hook Replacement

    gooseworks-ai/goose-skills

    Replace an existing video's opening with a supplied clip or free kinetic text hook while retaining and verifying every original body frame, audio, captions and ending.

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Blog Feed Monitor

    gooseworks-ai/goose-skills

    Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites.

    1.2k GitHub starsUsed in 1 repo~578 tokens
    Auto-check passed
  • Competitor Post Engagers

    gooseworks-ai/goose-skills

    Find leads by scraping engagers from a competitor's top LinkedIn posts.

    1.2k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Render Chatgpt Chat

    gooseworks-ai/goose-skills

    Assemble a ChatGPT chat-reveal video ad from a thread + timeline JSON — one continuous Playwright recording of a ChatGPT mobile chat (user types with the iOS keyboard up → taps send → keyboard…

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about Site Content Catalog

What does Site Content Catalog do?

Crawl a website's sitemap and blog index to build a complete content inventory. Site Content Catalog is an agent skill from gooseworks-ai/goose-skills. Crawl a website's sitemap and blog index to build a complete content inventory.

When should I use Site Content Catalog?

Site Content Catalog fits situations like: tasks that involve Content strategy; tasks that involve Web scraping; tasks that involve Brand voice and tone.

How do I install Site Content Catalog in Claude Code?

Run `npx skills add gooseworks-ai/goose-skills --skill site-content-catalog -a claude-code`. Or copy the skill folder (skills/seo/capabilities/site-content-catalog in gooseworks-ai/goose-skills) into .claude/skills/site-content-catalog in your project. Claude Code loads it when a task matches its description.

How do I install Site Content Catalog in Codex?

Run `npx skills add gooseworks-ai/goose-skills --skill site-content-catalog -a codex`. Or copy the skill folder (skills/seo/capabilities/site-content-catalog in gooseworks-ai/goose-skills) into .agents/skills/site-content-catalog in your project. Codex loads it when a task matches its description.

Can I use Site Content Catalog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gooseworks-ai/goose-skills --skill site-content-catalog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-content-catalog, .gemini/skills/site-content-catalog, .github/skills/site-content-catalog and .opencode/skills/site-content-catalog in your project.

What does Site Content Catalog need to run?

Going by SKILL.md and its folder, Site Content Catalog needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named APIFY_API_TOKEN. Our summary lists: Python 3; A credential in APIFY_API_TOKEN.

Does Site Content Catalog access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Site Content Catalog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Site Content Catalog use?

Site Content Catalog is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Site Content Catalog use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Site Content Catalog?

Skills that share tags, products or a category with Site Content Catalog: Caption Writer (stevenflanagan1/social-ai-team, 244 stars), SEO Coach (akseolabs-seo/seo-coach, 146 stars), E2E SEO Assistant (irinabuht12-oss/marketing-skills, 4.1k stars) and Brand Voice Content Creator (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Site Content Catalog?

gooseworks-ai (a GitHub organization) maintains it in gooseworks-ai/goose-skills, which has 1,240 GitHub stars. The repository holds 273 skills in this directory. The repository was last updated on October 8, 2026.

Source: gooseworks-ai/goose-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.