Firecrawl Page Scrape Integration
firecrawl/firecrawl
Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.
Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites.
$ npx skills add firecrawl/web-agent --skill structured-extraction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install firecrawl/web-agent structured-extraction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .claude/skills/structured-extraction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .claude/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extractionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add firecrawl/web-agent --skill structured-extraction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install firecrawl/web-agent structured-extraction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .agents/skills/structured-extraction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .agents/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add firecrawl/web-agent --skill structured-extraction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install firecrawl/web-agent structured-extraction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .cursor/skills/structured-extraction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .cursor/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/firecrawl/web-agent.git --path agent-core/src/skills/definitions/structured-extraction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add firecrawl/web-agent --skill structured-extraction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install firecrawl/web-agent structured-extraction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .gemini/skills/structured-extraction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .gemini/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install firecrawl/web-agent structured-extractionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add firecrawl/web-agent --skill structured-extraction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .github/skills/structured-extraction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .github/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add firecrawl/web-agent --skill structured-extraction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install firecrawl/web-agent structured-extraction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/firecrawl/web-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agent-core/src/skills/definitions/structured-extraction .opencode/skills/structured-extraction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "structured-extraction" agent skill from https://github.com/firecrawl/web-agent/tree/main/agent-core/src/skills/definitions/structured-extraction into .opencode/skills/structured-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "structured-extraction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
structured-extractionPlans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites.
The skill gives the agent a strategy for each kind of extraction task. A simple query means searching, scraping promising results with a targeted question, then building the object. Single-target research gathers several fields about one entity. For lists, the agent checks whether the list page already holds every requested detail and otherwise fetches item pages, using parallel workers through spawnAgents only when there are roughly five or more items or sources.
Crawling a whole site starts with sitemap.xml and robots.txt, then checks the entry page for pagination and categories, and uses interact to click through pages. Scraping advice is to prefer a scrape with a targeted query over raw page dumps, always ask about pagination, and never retry a 404 or bot check but move to other sources.
Output rules are strict: match the schema exactly, include every required field, use null for missing values, keep arrays as arrays and numbers as numbers (10.99, not a price string). Data from several sources can be merged with jq through bashExec. Before the result is passed to formatOutput, the agent checks required fields, types and duplicate array entries.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f023adf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Structured Web Data Extraction loads about 760 tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 415 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from firecrawl/web-agent at commit f023adf, republished under its MIT licence (© firecrawl). 415 words, ~760 tokens.
.claude/skills/structured-extraction/SKILL.md (or your agent's skills folder).Use this skill when extracting data that must match a specific JSON schema.
jq -s '.[0] * .[1]' /data/part1.json /data/part2.json > /data/merged.jsonBefore calling formatOutput, verify:
When you have gathered ALL data, call formatOutput with format "json" and the structured data. Do NOT stream data inline as markdown tables or JSON code blocks. Do NOT skip formatOutput — downstream systems depend on the structured output.
© firecrawl, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in agent-core/src/skills/definitions/structured-extraction of firecrawl/web-agent.
Open the folder on GitHubat commit f023adf
Structured Web Data Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Structured Web Data Extraction this skillfirecrawl/web-agent | 1.2k | — | ~760 | Automated safety check: Pass | MIT | |
| Firecrawl Page Scrape Integrationfirecrawl/firecrawl | 189k | 1 repos | ~944 | Automated safety check: Pass | ISC | |
| Firecrawl App Integrationfirecrawl/firecrawl | 189k | — | ~2.1k | Automated safety check: Notes | ISC | |
| Firecrawl Interact Integrationfirecrawl/firecrawl | 189k | 1 repos | ~731 | Automated safety check: Pass | ISC | |
| 9Router Web Fetchdecolua/9router | 30k | — | ~941 | Automated safety check: Pass | MIT | |
| Firecrawl Scrapefirecrawl/skills | 115 | — | ~1.8k | Automated safety check: Pass | ISC |
firecrawl/firecrawl
Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.
firecrawl/firecrawl
Adds web search, scraping, structured extraction and browser interaction to application code using Firecrawl's scrape, search and interact endpoints.
firecrawl/firecrawl
Guides adding Firecrawl's /interact endpoint to product code for pages that need clicks, forms, pagination or logged-in flows beyond plain scraping.
decolua/9router
Fetches a web page through a 9Router server's web fetch endpoint and returns it as markdown, plain text or HTML, using one of several extraction providers.
firecrawl/skills
Read a known webpage or execute a discovered workflow or data-provider capability.
firecrawl/skills
Autonomously navigate websites and extract structured data across pages.
firecrawl/web-agent
Compares two or more products or companies on pricing, features and positioning by scraping their sites, and returns a normalized JSON matrix.
firecrawl/web-agent
Pulls a public company's latest 10-K or 10-Q figures and analyst consensus from SEC EDGAR and Yahoo Finance, then cross-checks the two sources.
firecrawl/web-agent
Extracts every pricing tier from a SaaS, API, cloud or LLM vendor's pricing page and normalizes it into one structure, with optional price monitoring.
firecrawl/web-agent
Runs multi-source web research with a search plan, targeted extraction and cross-checking, rating confidence by how many sources agree and organizing results by subtopic.
firecrawl/web-agent
Playbook for scraping product names, prices, variants, stock and categories from online stores, including paginated listings and JavaScript-rendered storefronts.
Works with
Categories
Plans how to pull data from websites into an exact JSON schema, with separate approaches for single facts, one-entity research, item lists and whole sites. The skill gives the agent a strategy for each kind of extraction task. A simple query means searching, scraping promising results with a targeted question, then building the object.
Structured Web Data Extraction fits situations like: collecting product, company or listing data from websites into a fixed JSON schema; scraping a paginated catalogue where every item must appear exactly once; researching one entity across several pages and filling every field of a schema.
Run `npx skills add firecrawl/web-agent --skill structured-extraction -a claude-code`. Or copy the skill folder (agent-core/src/skills/definitions/structured-extraction in firecrawl/web-agent) into .claude/skills/structured-extraction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add firecrawl/web-agent --skill structured-extraction -a codex`. Or copy the skill folder (agent-core/src/skills/definitions/structured-extraction in firecrawl/web-agent) into .agents/skills/structured-extraction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add firecrawl/web-agent --skill structured-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/structured-extraction, .gemini/skills/structured-extraction, .github/skills/structured-extraction and .opencode/skills/structured-extraction in your project.
Going by SKILL.md and its folder, Structured Web Data Extraction needs the command-line tools its instructions call (jq). Our summary lists: Firecrawl web agent tools such as search, scrape, interact and formatOutput.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Structured Web Data Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 760 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Structured Web Data Extraction: Firecrawl Page Scrape Integration (firecrawl/firecrawl, 189k stars), Firecrawl App Integration (firecrawl/firecrawl, 189k stars), Firecrawl Interact Integration (firecrawl/firecrawl, 189k stars) and 9Router Web Fetch (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
firecrawl (a GitHub organization) maintains it in firecrawl/web-agent, which has 1,240 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 19, 2026.
Source: firecrawl/web-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.