Agent skill

Structured Web Data Extraction

by firecrawl in firecrawl/web-agent

Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites.

MITAuto-check passedData & Analytics

Install Structured Web Data Extraction

skills CLI
$ npx skills add firecrawl/web-agent --skill structured-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install firecrawl/web-agent structured-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .claude/skills/structured-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
structured-extraction
GitHub stars
1.2k
Token cost
~760 tokens
SKILL.md length
415 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites.

  • Works in 3 steps: Search for relevant results. → Scrape promising results with a targeted… → Build the result object and call…
  • Collecting product, company or listing data from websites into a fixed JSON schema
  • SKILL.md covers Strategy by task type, Scraping for structured data, Building the output and Validation before output, plus 1 more section
  • Calls jq

What it does

The skill gives the agent a strategy for each kind of extraction task. A simple query means searching, scraping promising results with a targeted question, then building the object. Single-target research gathers several fields about one entity. For lists, the agent checks whether the list page already holds every requested detail and otherwise fetches item pages, using parallel workers through spawnAgents only when there are roughly five or more items or sources.

Crawling a whole site starts with sitemap.xml and robots.txt, then checks the entry page for pagination and categories, and uses interact to click through pages. Scraping advice is to prefer a scrape with a targeted query over raw page dumps, always ask about pagination, and never retry a 404 or bot check but move to other sources.

Output rules are strict: match the schema exactly, include every required field, use null for missing values, keep arrays as arrays and numbers as numbers (10.99, not a price string). Data from several sources can be merged with jq through bashExec. Before the result is passed to formatOutput, the agent checks required fields, types and duplicate array entries.

When your agent uses it

  • Collecting product, company or listing data from websites into a fixed JSON schema
  • Scraping a paginated catalogue where every item must appear exactly once
  • Researching one entity across several pages and filling every field of a schema

Example prompts

  • “Extract the name, price and rating of every laptop on this store category into my product schema.”
  • “Pull founders, funding rounds and headquarters for this company into the schema in schema.json.”
  • “Crawl the whole blog and return each post's title, author and date as a JSON array.”

Requirements

  • Firecrawl web agent tools such as search, scrape, interact and formatOutput

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Search for relevant results.
  2. Scrape promising results with a targeted query.
  3. Build the result object and call formatOutput immediately.

What it can do on your machine

Read from SKILL.md and the folder at commit f023adf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Structured Web Data Extraction loads about 760 tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 415 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~760

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from firecrawl/web-agent at commit f023adf, republished under its MIT licence (© firecrawl). 415 words, ~760 tokens.

Download SKILL.mdSave it as .claude/skills/structured-extraction/SKILL.md (or your agent's skills folder).
name
structured-extraction
description
Extract structured data matching a JSON schema from websites. Handles complex nested schemas, arrays, pagination, and validation. Always outputs via formatOutput.
category
Research

Structured Extraction

Use this skill when extracting data that must match a specific JSON schema.

Strategy by task type

Simple query (single fact or small object)
  1. Search for relevant results.
  2. Scrape promising results with a targeted query.
  3. Build the result object and call formatOutput immediately.
Single target research (one entity, multiple fields)
  1. Search for relevant URLs.
  2. Scrape to extract data — stay in the orchestrator unless you have many independent sources (roughly 5+) where parallel workers clearly help.
  3. Compile findings and call formatOutput.
List of items (array in schema)
  1. Search/scrape to get the list of items.
  2. Are all requested details included in the list?
    • Yes: Build the result and call formatOutput.
    • No: If there are many items (roughly 5+), use spawnAgents so each worker gets the item and fields; otherwise fetch details sequentially in the orchestrator.
  3. Aggregate all results and call formatOutput.
All items from a website
  1. Check sitemaps (sitemap.xml, robots.txt) for an easy route to all pages.
  2. Scrape the entry page. Determine: pagination? Categories? Subcategories?
  3. For pagination, use interact to click through every page.
  4. For categories, scrape each category — use spawnAgents only when many independent categories warrant parallel fan-out.
  5. Aggregate and call formatOutput.

Scraping for structured data

  • PREFER scrape with a targeted query over raw page dumps. It keeps context lean.
  • When scraping lists, ALWAYS ask about pagination in your query: "How many total results? Is there a next page?"
  • For many independent URLs (roughly 5+), spawnAgents can help — each worker gets specific URLs and fields. Fewer URLs: handle in the orchestrator.
  • If a scrape returns a 404 or bot-check, do NOT retry. Move on to alternative sources.
Show full SKILL.md (140 more words)Show less

Building the output

  • Match the schema EXACTLY. Every required field must be present.
  • Use null for missing fields — never omit keys.
  • Arrays must be arrays even for single items.
  • Numbers must be actual numbers, not strings (10.99 not "$10.99").
  • Use bashExec with jq to merge data from multiple sources:
    jq -s '.[0] * .[1]' /data/part1.json /data/part2.json > /data/merged.json

Validation before output

Before calling formatOutput, verify:

  1. All required fields from the schema are present.
  2. Types match (numbers are numbers, arrays are arrays).
  3. No duplicate entries in arrays.
  4. Source URLs are included where the schema has citation fields.

CRITICAL: Always call formatOutput

When you have gathered ALL data, call formatOutput with format "json" and the structured data. Do NOT stream data inline as markdown tables or JSON code blocks. Do NOT skip formatOutput — downstream systems depend on the structured output.

© firecrawl, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in agent-core/src/skills/definitions/structured-extraction of firecrawl/web-agent.

Open the folder on GitHubat commit f023adf

Compare with similar skills

Structured Web Data Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Structured Web Data Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Structured Web Data Extraction this skillfirecrawl/web-agent1.2k—~760Automated safety check: PassMIT
Firecrawl Page Scrape Integrationfirecrawl/firecrawl189k1 repos~944Automated safety check: PassISC
Firecrawl App Integrationfirecrawl/firecrawl189k—~2.1kAutomated safety check: NotesISC
Firecrawl Interact Integrationfirecrawl/firecrawl189k1 repos~731Automated safety check: PassISC
9Router Web Fetchdecolua/9router30k—~941Automated safety check: PassMIT
Firecrawl Scrapefirecrawl/skills115—~1.8kAutomated safety check: PassISC

Similar skills

  • Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.

    189k GitHub starsUsed in 1 repo~944 tokens
    Data & AnalyticsAuto-check passed
  • Firecrawl App Integration

    firecrawl/firecrawl

    Adds web search, scraping, structured extraction and browser interaction to application code using Firecrawl's scrape, search and interact endpoints.

    189k GitHub stars~2.1k tokensUpdated today
    Data & AnalyticsAuto-check: notes
  • Guides adding Firecrawl's /interact endpoint to product code for pages that need clicks, forms, pagination or logged-in flows beyond plain scraping.

    189k GitHub starsUsed in 1 repo~731 tokens
    Data & AnalyticsAuto-check passed
  • 9Router Web Fetch

    decolua/9router

    Fetches a web page through a 9Router server's web fetch endpoint and returns it as markdown, plain text or HTML, using one of several extraction providers.

    30k GitHub stars~941 tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Firecrawl Scrape

    firecrawl/skills

    Read a known webpage or execute a discovered workflow or data-provider capability.

    115 GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Firecrawl Agent

    firecrawl/skills

    Autonomously navigate websites and extract structured data across pages.

    115 GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from firecrawl/web-agent

  • Competitor Comparison Matrix

    firecrawl/web-agent

    Compares two or more products or companies on pricing, features and positioning by scraping their sites, and returns a normalized JSON matrix.

    1.2k GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Financial Research

    firecrawl/web-agent

    Pulls a public company's latest 10-K or 10-Q figures and analyst consensus from SEC EDGAR and Yahoo Finance, then cross-checks the two sources.

    1.2k GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Vendor Pricing Tracker

    firecrawl/web-agent

    Extracts every pricing tier from a SaaS, API, cloud or LLM vendor's pricing page and normalizes it into one structure, with optional price monitoring.

    1.2k GitHub stars~1.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Deep Research

    firecrawl/web-agent

    Runs multi-source web research with a search plan, targeted extraction and cross-checking, rating confidence by how many sources agree and organizing results by subtopic.

    1.2k GitHub stars~278 tokensUpdated 5 mo ago
    Auto-check passed
  • Playbook for scraping product names, prices, variants, stock and categories from online stores, including paginated listings and JavaScript-rendered storefronts.

    1.2k GitHub stars~344 tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Structured Web Data Extraction

What does Structured Web Data Extraction do?

Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites. The skill gives the agent a strategy for each kind of extraction task. A simple query means searching, scraping promising results with a targeted question, then building the object.

When should I use Structured Web Data Extraction?

Structured Web Data Extraction fits situations like: collecting product, company or listing data from websites into a fixed JSON schema; scraping a paginated catalogue where every item must appear exactly once; researching one entity across several pages and filling every field of a schema.

How do I install Structured Web Data Extraction in Claude Code?

Run `npx skills add firecrawl/web-agent --skill structured-extraction -a claude-code`. Or copy the skill folder (agent-core/src/skills/definitions/structured-extraction in firecrawl/web-agent) into .claude/skills/structured-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Structured Web Data Extraction in Codex?

Run `npx skills add firecrawl/web-agent --skill structured-extraction -a codex`. Or copy the skill folder (agent-core/src/skills/definitions/structured-extraction in firecrawl/web-agent) into .agents/skills/structured-extraction in your project. Codex loads it when a task matches its description.

Can I use Structured Web Data Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add firecrawl/web-agent --skill structured-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/structured-extraction, .gemini/skills/structured-extraction, .github/skills/structured-extraction and .opencode/skills/structured-extraction in your project.

What does Structured Web Data Extraction need to run?

Going by SKILL.md and its folder, Structured Web Data Extraction needs the command-line tools its instructions call (jq). Our summary lists: Firecrawl web agent tools such as search, scrape, interact and formatOutput.

Does Structured Web Data Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Structured Web Data Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Structured Web Data Extraction use?

Structured Web Data Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Structured Web Data Extraction use?

About 760 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Structured Web Data Extraction?

Skills that share tags, products or a category with Structured Web Data Extraction: Firecrawl Page Scrape Integration (firecrawl/firecrawl, 189k stars), Firecrawl App Integration (firecrawl/firecrawl, 189k stars), Firecrawl Interact Integration (firecrawl/firecrawl, 189k stars) and 9Router Web Fetch (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Structured Web Data Extraction?

firecrawl (a GitHub organization) maintains it in firecrawl/web-agent, which has 1,240 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 19, 2026.

Source: firecrawl/web-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.