Agent skill

Firecrawl Site Crawl Extension

by AgriciDaniel in AgriciDaniel/claude-seo

Crawls, maps, and scrapes an entire site through the Firecrawl MCP server for broken-link checks and content inventories.

MITAuto-check passedMarketing & SEO

Install Firecrawl Site Crawl Extension

skills CLI
$ npx skills add AgriciDaniel/claude-seo --skill seo-firecrawl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AgriciDaniel/claude-seo seo-firecrawl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AgriciDaniel/claude-seo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/extensions/firecrawl/skills/seo-firecrawl .claude/skills/seo-firecrawl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
seo-firecrawl
GitHub stars
19k
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
907 words
Files
2
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

Crawls, maps, and scrapes an entire site through the Firecrawl MCP server for broken-link checks and content inventories.

  • Works in 5 steps: Comprehensive audit crawl: Crawl full… → Section-focused crawl: Use includePaths… → Broken link detection: Crawl with… → …
  • Discovering a site's full URL structure quickly and cheaply
  • SKILL.md covers Quick Reference, Commands, Cross-Skill Integration and Error Handling
  • Needs FIRECRAWL_API_KEY

What it does

Three Firecrawl MCP tools back the extension: firecrawl_map for a fast, content-free list of every URL, firecrawl_crawl for a full crawl with content and metadata that honors include and exclude path globs, and a per-page scrape that renders JavaScript for SPA content. A typical audit maps the site first to find its URLs cheaply, filters down to the most important pages, then crawls only those for full content.

When your agent uses it

  • Discovering a site's full URL structure quickly and cheaply
  • Finding broken links across an entire site
  • Auditing a JavaScript-rendered site that a plain fetch can't see

Example prompts

  • “Map this site's URL structure before we decide what to crawl.”
  • “Crawl the /blog section and check every link for 404s.”
  • “Scrape this SPA page with JS rendering enabled.”

Requirements

  • Firecrawl MCP server
  • Compatibility (from SKILL.md): Requires Firecrawl MCP server

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Comprehensive audit crawl: Crawl full site, extract all pages for subagent analysis
  2. Section-focused crawl: Use includePaths to audit only /blog/* or /products/*
  3. Broken link detection: Crawl with ["links"] format, check all hrefs for 404s
  4. Content inventory: Extract all page titles, meta descriptions, H1s at scale
  5. SPA/JS-rendered sites: Firecrawl renders JavaScript, solving the Issue #11 problem

What it can do on your machine

Read from SKILL.md and the folder at commit 4b99de2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FIRECRAWL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Firecrawl MCP server

    From compatibility in the SKILL.md frontmatter.

Context cost

Firecrawl Site Crawl Extension loads about 2k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 907 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AgriciDaniel/claude-seo at commit 4b99de2, republished under its MIT licence (© AgriciDaniel). 907 words, ~2,020 tokens.

Download SKILL.mdSave it as .claude/skills/seo-firecrawl/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
seo-firecrawl
description
Full-site crawling, scraping, and site mapping via Firecrawl MCP. Use when user says "crawl site", "map site", "full crawl", "find all pages", "broken links", "site structure", "discover pages", "JS rendering", or needs site-wide analysis.
compatibility
Requires Firecrawl MCP server
user-invocable
true
argument-hint
[command] <url>
license
MIT
metadata.author
AgriciDaniel
metadata.version
2.4.2
metadata.category
seo

Firecrawl Extension for Claude SEO

This skill requires the Firecrawl extension to be installed:

bash
./extensions/firecrawl/install.sh

Check availability: Before using any Firecrawl tool, verify the MCP server is connected by checking if firecrawl_scrape or any Firecrawl tool is available. If tools are not available, inform the user the extension is not installed and provide install instructions.

Quick Reference

CommandPurpose
/seo firecrawl crawl <url>Full-site crawl with content extraction
/seo firecrawl map <url>Discover site structure (URLs only, fast)
/seo firecrawl scrape <url>Single-page scrape with JS rendering
/seo firecrawl search <query> <url>Search within a crawled site

Commands

crawl -- Full-Site Crawl

Crawl an entire website starting from the given URL. Returns page content, metadata, and links for all discovered pages.

MCP Tool: firecrawl_crawl

Parameters:

  • url (required): Starting URL to crawl
  • limit: Max pages to crawl (default: 100, max: 500)
  • maxDepth: Max link depth from start URL (default: 3)
  • includePaths: Array of glob patterns to include (e.g., ["/blog/*"])
  • excludePaths: Array of glob patterns to exclude (e.g., ["/admin/*", "/api/*"])
  • scrapeOptions.formats: Output formats -- ["markdown", "html", "links"]

SEO Usage Patterns:

  1. Comprehensive audit crawl: Crawl full site, extract all pages for subagent analysis
  2. Section-focused crawl: Use includePaths to audit only /blog/* or /products/*
  3. Broken link detection: Crawl with ["links"] format, check all hrefs for 404s
  4. Content inventory: Extract all page titles, meta descriptions, H1s at scale
  5. SPA/JS-rendered sites: Firecrawl renders JavaScript, solving the Issue #11 problem

Example orchestration for /seo audit:

1. firecrawl_map(url) -> get all URLs (fast, no content)
2. Filter to top 50 most important pages (homepage, key sections)
3. firecrawl_crawl(url, limit=50) -> get full content
4. Feed content to seo-technical, seo-content, seo-schema agents

Cost awareness:

  • Free tier: 500 credits/month
  • 1 credit = 1 page crawled or scraped
  • Map operations are cheaper (0.5 credits per URL discovered)
  • Always inform user of estimated credit usage before large crawls
map -- Site Structure Discovery

Discover all URLs on a website without fetching content. Fast and credit-efficient.

MCP Tool: firecrawl_map

Parameters:

  • url (required): Website URL to map
  • limit: Max URLs to discover (default: 5000)
  • search: Optional search term to filter URLs

SEO Usage Patterns:

  1. Sitemap comparison: Map site, compare discovered URLs vs XML sitemap
  2. Orphan page detection: URLs in sitemap but not linked from any page
  3. Crawl budget analysis: Total indexable pages vs pages linked from homepage
  4. URL pattern analysis: Identify URL structure patterns, duplicates, parameter bloat
  5. Pre-audit discovery: Run map first, then targeted crawl on key sections

Output: Array of URLs. Present as:

Site: example.com
Pages discovered: 342

URL Pattern Breakdown:
  /blog/*          - 128 pages (37%)
  /products/*      - 89 pages (26%)
  /category/*      - 45 pages (13%)
  /pages/*         - 32 pages (9%)
  / (root pages)   - 48 pages (14%)
scrape -- Single-Page Deep Scrape

Scrape a single page with full JavaScript rendering. More thorough than fetch_page.py because it executes JS and waits for dynamic content.

MCP Tool: firecrawl_scrape

Parameters:

  • url (required): Page URL to scrape
  • formats: Output formats -- ["markdown", "html", "links", "screenshot"]
  • onlyMainContent: Strip nav/footer/sidebar (default: true)
  • waitFor: CSS selector or milliseconds to wait for content
  • timeout: Request timeout in ms (default: 30000)
  • actions: Browser actions before scraping (click, scroll, wait)

SEO Usage Patterns:

  1. SPA content extraction: Scrape JS-rendered React/Vue/Angular pages
  2. Dynamic content audit: Pages with lazy-loaded content below the fold
  3. Paywall/login detection: Identify content behind authentication walls
  4. Main content extraction: Use onlyMainContent for clean E-E-A-T analysis
  5. Screenshot capture: Use screenshot format for visual analysis

When to use scrape vs fetch_page.py:

ScenarioUse
Static HTML pagefetch_page.py (no API cost)
JS-rendered SPAfirecrawl_scrape (renders JS)
Need response headersfetch_page.py (returns headers)
Need clean markdownfirecrawl_scrape (better extraction)
Rate-limited/blockedfirecrawl_scrape (handles anti-bot)
Show full SKILL.md (371 more words)Show less

Search within a website for specific content. Useful for finding pages related to a topic without crawling everything.

MCP Tool: firecrawl_search

Parameters:

  • query (required): Search query
  • url (required): Website to search within
  • limit: Max results (default: 10)
  • scrapeOptions.formats: Output format for matched pages

SEO Usage Patterns:

  1. Content gap validation: Search for a keyword on the site to check if content exists
  2. Internal linking opportunities: Find pages mentioning a topic that could link to each other
  3. Duplicate content detection: Search for key phrases to find near-duplicates
  4. Competitor content research: Search competitor site for specific topics

Cross-Skill Integration

With seo-audit (full audit)

When Firecrawl is available during /seo audit:

  1. Use firecrawl_map to discover all site URLs
  2. Compare with XML sitemap (seo-sitemap) to find orphan/missing pages
  3. Select top pages for deep analysis
  4. Feed crawled content to all subagents (technical, content, schema, geo)
  5. Report total crawlable pages, URL patterns, and crawl depth
With seo-technical
  • Broken link detection: crawl all internal links, check for 404s
  • Redirect chain mapping: follow all redirects, flag chains > 2 hops
  • Mixed content detection: check HTTP resources on HTTPS pages
  • Canonical verification: compare canonical URLs with actual URLs
With seo-sitemap
  • Sitemap coverage: % of crawled pages present in sitemap
  • Orphan pages: pages found by crawl but missing from sitemap
  • Stale sitemap entries: URLs in sitemap that return 404/410
With seo-content
  • Content extraction: feed clean markdown to E-E-A-T analysis
  • Thin content detection: identify pages with < 300 words at scale
  • Duplicate content: compare content across pages for near-duplicates
With seo-schema
  • Schema extraction: pull JSON-LD from all crawled pages
  • Schema coverage: % of pages with structured data
  • Schema validation: batch-validate extracted schemas

Error Handling

ErrorCauseResolution
FIRECRAWL_API_KEY not setMCP not configuredRun ./extensions/firecrawl/install.sh
402 Payment RequiredCredits exhaustedCheck usage at firecrawl.dev/app, upgrade plan
429 Too Many RequestsRate limitedWait 60s, reduce crawl concurrency
408 TimeoutPage too slow to renderIncrease timeout, try without JS rendering
403 ForbiddenSite blocks crawlingCheck robots.txt, may need to skip this site

Graceful fallback: If Firecrawl is unavailable, inform the user and suggest:

  1. Use fetch_page.py for single-page analysis (no API cost)
  2. Use WebFetch tool for basic HTML retrieval
  3. Install Firecrawl: ./extensions/firecrawl/install.sh

© AgriciDaniel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in extensions/firecrawl/skills/seo-firecrawl of AgriciDaniel/claude-seo.

  • SKILL.md
  • LICENSE.txt

Open the folder on GitHubat commit 4b99de2

Used in 2 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in AgriciDaniel/claude-seo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Firecrawl Site Crawl Extension next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Firecrawl Site Crawl Extension compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Firecrawl Site Crawl Extension this skillAgriciDaniel/claude-seo19k1 repos~2kAutomated safety check: PassMIT
SEO Sitemapseranking/seo-skills160—~2.3kAutomated safety check: PassMIT
SEO Firecrawlseranking/seo-skills160—~2.3kAutomated safety check: PassMIT
SEO Imagesseranking/seo-skills160—~6kAutomated safety check: PassMIT
Firecrawl Scrapeparcadei/Continuous-Claude-v33.9k1 repos~254Automated safety check: NotesMIT
Firecrawl Developer Indexfirecrawl/skills117—~2kAutomated safety check: PassISC

Similar skills

  • SEO Sitemap

    seranking/seo-skills

    Pull a domain's XML sitemap (and sitemap-of-sitemaps), then compare against the most recent SE Ranking website audit.

    160 GitHub stars~2.3k tokensUpdated 3 mo ago
    Marketing & SEOAuto-check passed
  • SEO Firecrawl

    seranking/seo-skills

    Ad-hoc web scraping, site mapping, and full-site crawling via Firecrawl MCP.

    160 GitHub stars~2.3k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • SEO Images

    seranking/seo-skills

    Image SEO audit for a URL or domain. An agent skill from seranking/seo-skills.

    160 GitHub stars~6k tokensUpdated 3 mo ago
    Marketing & SEOAuto-check passed
  • Firecrawl Scrape

    parcadei/Continuous-Claude-v3

    Scrape web pages and extract content via Firecrawl MCP. An agent skill from parcadei/Continuous-Claude-v3.

    3.9k GitHub starsUsed in 1 repo~254 tokens
    Data & AnalyticsAuto-check: notes
  • Search an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites.

    117 GitHub stars~2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Serving The API

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when the user wants a long-running HTTP service for scrape/crawl/map instead of one-shot CLI calls or the MCP server — for example wiring crawlberg into other apps over REST.

    1.3k GitHub stars~1.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from AgriciDaniel/claude-seo

All 33 skills in this repo
  • Hreflang and International SEO

    AgriciDaniel/claude-seo

    Audits, validates and generates hreflang tags for multi-language and multi-region sites in HTML, HTTP headers or XML sitemaps, flagging common code and return-tag mistakes.

    19k GitHub starsUsed in 5 repos~3.4k tokens
    Auto-check passed
  • Google SEO APIs

    AgriciDaniel/claude-seo

    Pulls real Google data for SEO work: Search Console, PageSpeed Insights, CrUX field data, the Indexing API and GA4 organic traffic, through /seo google commands.

    19k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • SEO Keyword Clustering

    AgriciDaniel/claude-seo

    Clusters keywords by how much their search results overlap and designs a hub-and-spoke content plan with an internal link matrix and an interactive cluster map.

    19k GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • SEO Content Brief Generator

    AgriciDaniel/claude-seo

    Builds research-backed SEO content briefs with competitor scoring, per-section word counts and page-type templates, for new pages or improving existing ones.

    19k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed
  • FLOW SEO Framework

    AgriciDaniel/claude-seo

    Brings the FLOW framework's stage-specific SEO prompts into the agent, from keyword discovery through backlinks, on-page work and conversion to local SEO, loaded on demand.

    19k GitHub starsUsed in 2 repos~1.4k tokens
    Auto-check passed
  • SEO Image Generator

    AgriciDaniel/claude-seo

    Generates Open Graph previews, blog hero images, product photos and infographics for SEO use through Gemini image tools and the banana extension.

    19k GitHub starsUsed in 2 repos~2.1k tokens
    Auto-check passed

Questions about Firecrawl Site Crawl Extension

What does Firecrawl Site Crawl Extension do?

Crawls, maps, and scrapes an entire site through the Firecrawl MCP server for broken-link checks and content inventories. Three Firecrawl MCP tools back the extension: firecrawl_map for a fast, content-free list of every URL, firecrawl_crawl for a full crawl with content and metadata that honors include and exclude path globs, and a per-page scrape that renders JavaScript for SPA content. A typical audit maps the site first to find its URLs cheaply, filters down to the most important pages, then crawls only those for full content.

When should I use Firecrawl Site Crawl Extension?

Firecrawl Site Crawl Extension fits situations like: discovering a site's full URL structure quickly and cheaply; finding broken links across an entire site; auditing a JavaScript-rendered site that a plain fetch can't see.

How do I install Firecrawl Site Crawl Extension in Claude Code?

Run `npx skills add AgriciDaniel/claude-seo --skill seo-firecrawl -a claude-code`. Or copy the skill folder (extensions/firecrawl/skills/seo-firecrawl in AgriciDaniel/claude-seo) into .claude/skills/seo-firecrawl in your project. Claude Code loads it when a task matches its description.

How do I install Firecrawl Site Crawl Extension in Codex?

Run `npx skills add AgriciDaniel/claude-seo --skill seo-firecrawl -a codex`. Or copy the skill folder (extensions/firecrawl/skills/seo-firecrawl in AgriciDaniel/claude-seo) into .agents/skills/seo-firecrawl in your project. Codex loads it when a task matches its description.

Can I use Firecrawl Site Crawl Extension in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgriciDaniel/claude-seo --skill seo-firecrawl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/seo-firecrawl, .gemini/skills/seo-firecrawl, .github/skills/seo-firecrawl and .opencode/skills/seo-firecrawl in your project.

What does Firecrawl Site Crawl Extension need to run?

Going by SKILL.md and its folder, Firecrawl Site Crawl Extension needs credentials named FIRECRAWL_API_KEY. Our summary lists: Firecrawl MCP server. Compatibility (from SKILL.md): Requires Firecrawl MCP server.

Does Firecrawl Site Crawl Extension access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Firecrawl Site Crawl Extension safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Firecrawl Site Crawl Extension use?

Firecrawl Site Crawl Extension is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Firecrawl Site Crawl Extension use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Firecrawl Site Crawl Extension?

Skills that share tags, products or a category with Firecrawl Site Crawl Extension: SEO Sitemap (seranking/seo-skills, 160 stars), SEO Firecrawl (seranking/seo-skills, 160 stars), SEO Images (seranking/seo-skills, 160 stars) and Firecrawl Scrape (parcadei/Continuous-Claude-v3, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Firecrawl Site Crawl Extension?

AgriciDaniel (a GitHub user) maintains it in AgriciDaniel/claude-seo, which has 18,574 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 4, 2026.

Source: AgriciDaniel/claude-seo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.