Agent skill

Firecrawl Reference Architecture

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion.

MITAuto-check passedData & Analytics

Install Firecrawl Reference Architecture

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill firecrawl-reference-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace firecrawl-reference-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/firecrawl-reference-architecture .claude/skills/firecrawl-reference-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
firecrawl-reference-architecture
GitHub stars
2.8k
Token cost
~1.1k tokens
SKILL.md length
477 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion.

  • Works in 7 steps: Define tenants, source authorization,… → Place authentication, tenant isolation,… → Separate discovery into map/search,… → …
  • Building a reusable platform
  • SKILL.md covers Overview, Prerequisites, Current Contract and Authentication, plus 7 more sections
  • Needs FIRECRAWL_API_KEY

What it does

Firecrawl Reference Architecture is an agent skill from jeremylongshore/tons-of-skills-marketplace. Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion. Use when building a reusable platform. Trigger with "Firecrawl reference architecture", "Firecrawl ingestion system", or "Firecrawl RAG pipeline".

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/official-docs.md`). Compatibility notes: Designed for Claude Code; Firecrawl Cloud work requires network access

It sits in Data & Analytics, covering Web scraping. It works with Firecrawl. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Building a reusable platform
  • With Firecrawl reference architecture
  • Firecrawl ingestion system
  • Firecrawl RAG pipeline

Example prompts

  • “Firecrawl reference architecture”
  • “Firecrawl ingestion system”
  • “Firecrawl RAG pipeline”
  • “/firecrawl-reference-architecture”

Requirements

  • A credential in FIRECRAWL_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code; Firecrawl Cloud work requires network access
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Write, Edit

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Define tenants, source authorization, freshness, scale, formats, data classification, SLO/RPO/RTO, retention, deletion, and cost…
  2. Place authentication, tenant isolation, URL canonicalization, policy versioning, endpoint/format allowlists, budgets, and idempotency at…
  3. Separate discovery into map/search, retrieval into scrape/batch/crawl/parse, and model extraction into a validated stage. Use durable job…
  4. Persist normalized provenance and job/page state before downstream writes. Complete pagination and deduplicate by tenant, canonical…
  5. Run schema, origin-status, content-quality, malware/active-content, and prompt-injection gates before storage or agent context.
  6. Use an outbox or equivalent for indexing and publication; make retries idempotent and support tombstones across raw, normalized…
  7. Observe queue, job, origin, validation, freshness, spend, webhook, and deletion SLIs; test restore, reconciliation, cancellation, and…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FIRECRAWL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code; Firecrawl Cloud work requires network access

    From compatibility in the SKILL.md frontmatter.

Context cost

Firecrawl Reference Architecture loads about 1.1k tokens when it runs, and up to ~1.6k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 477 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 477 words, ~1,142 tokens.

Download SKILL.mdSave it as .claude/skills/firecrawl-reference-architecture/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
firecrawl-reference-architecture
description
Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion. Use when building a reusable platform. Trigger with "Firecrawl reference architecture", "Firecrawl ingestion system", or "Firecrawl RAG pipeline".
allowed-tools
Read, Glob, Grep, Write, Edit
compatibility
Designed for Claude Code; Firecrawl Cloud work requires network access
argument-hint
<repository-path> <workload-profile>
version
1.12.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, firecrawl, architecture, ingestion
model
inherit
effort
high

Firecrawl Governed Ingestion Architecture

Overview

Create explicit trust boundaries from request intake through deletion. Keep source discovery, content retrieval, validation, and downstream publication independently retryable and auditable.

Prerequisites

  • The target repository or integration path and the requested operator outcome.
  • The source authorization, data classification, and environment policy.
  • Current Firecrawl documentation, credentials only when needed, and an owner for approvals.

Current Contract

Firecrawl v2 provides scrape, crawl, map, search, batch scrape, parse, JSON extraction, agentic, and browser surfaces with different async, credit, retention, and availability contracts. The SDK can auto-wait and paginate, while explicit start/status methods support durable orchestration. Provider webhooks supplement but do not replace reconciliation.

Authentication

For authenticated Cloud operations, inject FIRECRAWL_API_KEY from an approved secret manager. REST requests use Authorization: Bearer with the key. Never print, commit, transmit, or place a key in a URL. Keyless access is suitable only where the current documentation explicitly allows it and the workload accepts its limits; production workflows should make identity and team ownership explicit.

Instructions

  1. Define tenants, source authorization, freshness, scale, formats, data classification, SLO/RPO/RTO, retention, deletion, and cost constraints.
  2. Place authentication, tenant isolation, URL canonicalization, policy versioning, endpoint/format allowlists, budgets, and idempotency at the intake gateway.
  3. Separate discovery into map/search, retrieval into scrape/batch/crawl/parse, and model extraction into a validated stage. Use durable job records for asynchronous work.
  4. Persist normalized provenance and job/page state before downstream writes. Complete pagination and deduplicate by tenant, canonical source, version, and content hash.
  5. Run schema, origin-status, content-quality, malware/active-content, and prompt-injection gates before storage or agent context.
  6. Use an outbox or equivalent for indexing and publication; make retries idempotent and support tombstones across raw, normalized, embedding, cache, and search stores.
  7. Observe queue, job, origin, validation, freshness, spend, webhook, and deletion SLIs; test restore, reconciliation, cancellation, and regional/provider failure.
Show full SKILL.md (177 more words)Show less

Tool Discipline

Use Read, Glob, and Grep to inspect code, configuration, tests, and evidence. Use Write/Edit only for approved implementation or documentation changes. Do not call Firecrawl, rotate keys, change account settings, scrape a target, or deploy merely because this skill was invoked.

Approval Boundaries

Require architecture and security/data approval before adding a new endpoint class, authenticated source, model extraction, long-term store, cross-region flow, or self-hosted provider.

Output

Return components and trust boundaries, sequence and state model, contracts, capacity and cost assumptions, security/privacy controls, failure and reconciliation paths, SLOs, tests, and phased rollout.

Error Handling

  • A stage lacks durable identity or idempotency: block asynchronous rollout.
  • Provider completion and stored counts diverge: reconcile pagination and downstream receipts before publication.
  • Deletion cannot propagate to derived stores: fail the data architecture review.

Examples

  • "Build a docs RAG pipeline" produces policy, acquisition, validation, outbox, index, and deletion stages.
  • "Call crawl directly from the browser" is replaced with a server-side governed gateway.

Resources

Read official Firecrawl evidence before relying on an endpoint, SDK method, plan limit, price, retention option, or self-hosted release.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/firecrawl-reference-architecture of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/official-docs.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Firecrawl Reference Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Firecrawl Reference Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Firecrawl Reference Architecture this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Firecrawl Scrapefirecrawl/skills117—~1.8kAutomated safety check: PassISC
Firecrawl Agentfirecrawl/skills117—~1.2kAutomated safety check: PassISC
Firecrawl Company Directoriesfirecrawl/skills117—~557Automated safety check: PassISC
Firecrawl Competitive Intelfirecrawl/skills117—~604Automated safety check: PassISC
Firecrawlzapier/connectors177—~4.2kAutomated safety check: PassElastic-2.0

Similar skills

  • Firecrawl Scrape

    firecrawl/skills

    Read a known webpage or execute a discovered workflow or data-provider capability.

    117 GitHub stars~1.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Firecrawl Agent

    firecrawl/skills

    Autonomously navigate websites and extract structured data across pages.

    117 GitHub stars~1.2k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Extract structured company lists from directories with Firecrawl.

    117 GitHub stars~557 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Monitor competitor pricing, features, changelogs, dashboards, and product changes with Firecrawl.

    117 GitHub stars~604 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Firecrawl

    zapier/connectors

    Official

    Agent-callable Firecrawl tools — scrape a URL to clean Markdown, crawl a site, search the web, map site URLs, and extract structured data.

    177 GitHub stars~4.2k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • SEO Firecrawl

    seranking/seo-skills

    Ad-hoc web scraping, site mapping, and full-site crawling via Firecrawl MCP.

    161 GitHub stars~2.3k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Firecrawl Reference Architecture

What does Firecrawl Reference Architecture do?

Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion. Firecrawl Reference Architecture is an agent skill from jeremylongshore/tons-of-skills-marketplace. Design a production Firecrawl v2 ingestion system with policy gateway, bounded acquisition, validation, provenance, storage, indexing, observability, and deletion.

When should I use Firecrawl Reference Architecture?

Firecrawl Reference Architecture fits situations like: building a reusable platform; with Firecrawl reference architecture; firecrawl ingestion system; firecrawl RAG pipeline.

How do I install Firecrawl Reference Architecture in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill firecrawl-reference-architecture -a claude-code`. Or copy the skill folder (skills/.curated/firecrawl-reference-architecture in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/firecrawl-reference-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Firecrawl Reference Architecture in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill firecrawl-reference-architecture -a codex`. Or copy the skill folder (skills/.curated/firecrawl-reference-architecture in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/firecrawl-reference-architecture in your project. Codex loads it when a task matches its description.

Can I use Firecrawl Reference Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill firecrawl-reference-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/firecrawl-reference-architecture, .gemini/skills/firecrawl-reference-architecture, .github/skills/firecrawl-reference-architecture and .opencode/skills/firecrawl-reference-architecture in your project.

What does Firecrawl Reference Architecture need to run?

Going by SKILL.md and its folder, Firecrawl Reference Architecture needs credentials named FIRECRAWL_API_KEY. Our summary lists: A credential in FIRECRAWL_API_KEY. Its frontmatter pre-approves these tools: Read, Glob, Grep, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code; Firecrawl Cloud work requires network access.

Does Firecrawl Reference Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Firecrawl Reference Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Firecrawl Reference Architecture use?

Firecrawl Reference Architecture is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Firecrawl Reference Architecture use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 454 tokens, read only when the agent opens those files.

What are the alternatives to Firecrawl Reference Architecture?

Skills that share tags, products or a category with Firecrawl Reference Architecture: Firecrawl Scrape (firecrawl/skills, 117 stars), Firecrawl Agent (firecrawl/skills, 117 stars), Firecrawl Company Directories (firecrawl/skills, 117 stars) and Firecrawl Competitive Intel (firecrawl/skills, 117 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Firecrawl Reference Architecture?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.