Agent skill

Cf Crawl

by davila7 in davila7/claude-code-templates

Crawl entire websites using Cloudflare Browser Rendering /crawl API.

MITAuto-check: notesKnowledge Management

Install Cf Crawl

skills CLI
$ npx skills add davila7/claude-code-templates --skill cf-crawl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates cf-crawl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/utilities/cf-crawl .claude/skills/cf-crawl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cf-crawl
GitHub stars
32k
Token cost
~2.6k tokens
SKILL.md length
766 words
Files
1
Skills in repo
478
Repo updated
First seen
Licence
MIT

At a glance

Crawl entire websites using Cloudflare Browser Rendering /crawl API.

  • Works in 6 steps: Load Credentials → Validate Credentials → Initiate Crawl → …
  • Tasks that involve Web scraping
  • SKILL.md covers Prerequisites, Workflow, Parameter Reference and Usage Examples, plus 2 more sections
  • Calls curl, python3 and cursor; reaches api.cloudflare.com; needs CLOUDFLARE_API_TOKEN and API_TOKEN

What it does

Cf Crawl is an agent skill from davila7/claude-code-templates. Crawl entire websites using Cloudflare Browser Rendering /crawl API. Initiates async crawl jobs, polls for completion, and saves results as markdown files. Useful for ingesting documentation sites, knowledge bases, or any web content into your project context. Requires CLOUDFLAREACCOUNTID and CLOUDFLAREAPITOKEN environment variables.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Knowledge Management, covering Web scraping, Static sites and blogs and Knowledge bases. It works with Cloudflare. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve Static sites and blogs
  • Tasks that involve Knowledge bases

Example prompts

  • “/cf-crawl”

Requirements

  • Python 3
  • A credential in CLOUDFLARE_API_TOKEN
  • A credential in API_TOKEN

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load Credentials
  2. Validate Credentials
  3. Initiate Crawl
  4. Poll for Completion
  5. Retrieve Results
  6. Save Results

What it can do on your machine

Read from SKILL.md and the folder at commit 46b4d8b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • python3
    • cursor

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.cloudflare.com

    Also links to:

    • dash.cloudflare.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CLOUDFLARE_API_TOKEN
    • API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cf Crawl loads about 2.6k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 766 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:25
    2. **Project `.env` file** - Read `.env` in the current working directory and extract the values
  • NoteMentions a .env fileSKILL.md:26
    3. **Project `.env.local` file** - Read `.env.local` in the current working directory
  • NoteMentions a .env fileSKILL.md:27
    4. **Home directory `.env`** - Read `~/.env` as a last resort
  • NoteMentions a .env fileSKILL.md:29
    To load from a `.env` file, parse it line by line looking for `CLOUDFLARE_ACCOUNT_ID=` and `CLOUDFLARE_API_TOKEN=` entri
  • NoteMentions a .env fileSKILL.md:32
    # Load from .env if vars are not already set
  • NoteMentions a .env fileSKILL.md:34
    for envfile in .env .env.local "$HOME/.env"; do
  • NoteMentions a .env fileSKILL.md:42
    l the user to add them to their project `.env` file:

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 46b4d8b, republished under its MIT licence (© davila7). 766 words, ~2,637 tokens.

Download SKILL.mdSave it as .claude/skills/cf-crawl/SKILL.md (or your agent's skills folder).
name
cf-crawl
description
Crawl entire websites using Cloudflare Browser Rendering /crawl API. Initiates async crawl jobs, polls for completion, and saves results as markdown files. Useful for ingesting documentation sites, knowledge bases, or any web content into your project context. Requires CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN environment variables.

Cloudflare Website Crawler

You are a web crawling assistant that uses Cloudflare's Browser Rendering /crawl REST API to crawl websites and save their content as markdown files for local use.

Prerequisites

The user must have:

  1. A Cloudflare account with Browser Rendering enabled
  2. CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN available (see below)

Workflow

When the user asks to crawl a website, follow this exact workflow:

Step 1: Load Credentials

Look for CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN in this order:

  1. Current environment variables - Check if already exported in the shell
  2. Project .env file - Read .env in the current working directory and extract the values
  3. Project .env.local file - Read .env.local in the current working directory
  4. Home directory .env - Read ~/.env as a last resort

To load from a .env file, parse it line by line looking for CLOUDFLARE_ACCOUNT_ID= and CLOUDFLARE_API_TOKEN= entries. Use this bash approach:

bash
# Load from .env if vars are not already set
if [ -z "$CLOUDFLARE_ACCOUNT_ID" ] || [ -z "$CLOUDFLARE_API_TOKEN" ]; then
  for envfile in .env .env.local "$HOME/.env"; do
    if [ -f "$envfile" ]; then
      eval "$(grep -E '^CLOUDFLARE_(ACCOUNT_ID|API_TOKEN)=' "$envfile" | sed 's/^/export /')"
    fi
  done
fi

If credentials are still missing after checking all sources, tell the user to add them to their project .env file:

CLOUDFLARE_ACCOUNT_ID=your-account-id
CLOUDFLARE_API_TOKEN=your-api-token

The API token needs "Browser Rendering - Edit" permission. Create one at Cloudflare Dashboard > API Tokens.

Step 2: Validate Credentials

Verify both variables are set and non-empty before proceeding.

Step 3: Initiate Crawl

Send a POST request to start the crawl job. Choose parameters based on user needs:

bash
curl -s -X POST "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "<TARGET_URL>",
    "limit": <NUMBER_OF_PAGES>,
    "formats": ["markdown"],
    "options": {
      "excludePatterns": ["**/changelog/**", "**/api-reference/**"]
    }
  }'

For incremental crawls, add the modifiedSince parameter (Unix timestamp in seconds):

bash
curl -s -X POST "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "<TARGET_URL>",
    "limit": <NUMBER_OF_PAGES>,
    "formats": ["markdown"],
    "modifiedSince": <UNIX_TIMESTAMP>
  }'

When --since is provided, convert to Unix timestamp: date -d "2026-03-10" +%s (Linux) or date -j -f "%Y-%m-%d" "2026-03-10" +%s (macOS).

The response returns a job ID:

json
{"success": true, "result": "job-uuid-here"}
Step 4: Poll for Completion

Poll the job status every 5 seconds until it completes:

bash
curl -s -X GET "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl/<JOB_ID>?limit=1" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" | python3 -c "import sys,json; d=json.load(sys.stdin); print(f'Status: {d[\"result\"][\"status\"]} | Finished: {d[\"result\"][\"finished\"]}/{d[\"result\"][\"total\"]}')"

Possible job statuses:

  • running - Still in progress, keep polling
  • completed - All pages processed
  • cancelled_due_to_timeout - Exceeded 7-day limit
  • cancelled_due_to_limits - Hit account limits
  • errored - Something went wrong
Step 5: Retrieve Results

When using modifiedSince, check for skipped pages to see what was unchanged:

bash
# See which pages were skipped (not modified since the given timestamp)
curl -s -X GET "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl/<JOB_ID>?status=skipped&limit=50" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}"

Fetch all completed records using pagination (cursor-based):

bash
curl -s -X GET "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl/<JOB_ID>?status=completed&limit=50" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}"

If there are more records, use the cursor value from the response:

bash
curl -s -X GET "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/browser-rendering/crawl/<JOB_ID>?status=completed&limit=50&cursor=<CURSOR>" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}"
Step 6: Save Results

Save each page's markdown content to a local directory. Use a script like:

bash
# Create output directory
mkdir -p .crawl-output

# Fetch and save all pages
python3 -c "
import json, os, re, sys, urllib.request

account_id = os.environ['CLOUDFLARE_ACCOUNT_ID']
api_token = os.environ['CLOUDFLARE_API_TOKEN']
job_id = '<JOB_ID>'
base = f'https://api.cloudflare.com/client/v4/accounts/{account_id}/browser-rendering/crawl/{job_id}'
outdir = '.crawl-output'
os.makedirs(outdir, exist_ok=True)

cursor = None
total_saved = 0

while True:
    url = f'{base}?status=completed&limit=50'
    if cursor:
        url += f'&cursor={cursor}'

    req = urllib.request.Request(url, headers={
        'Authorization': f'Bearer {api_token}'
    })
    with urllib.request.urlopen(req) as resp:
        data = json.load(resp)

    records = data.get('result', {}).get('records', [])
    if not records:
        break

    for rec in records:
        page_url = rec.get('url', '')
        md = rec.get('markdown', '')
        if not md:
            continue
        # Convert URL to filename
        name = re.sub(r'https?://', '', page_url)
        name = re.sub(r'[^a-zA-Z0-9]', '_', name).strip('_')[:120]
        filepath = os.path.join(outdir, f'{name}.md')
        with open(filepath, 'w') as f:
            f.write(f'<!-- Source: {page_url} -->\n\n')
            f.write(md)
        total_saved += 1

    cursor = data.get('result', {}).get('cursor')
    if cursor is None:
        break

print(f'Saved {total_saved} pages to {outdir}/')
"

Parameter Reference

Core Parameters
ParameterTypeDefaultDescription
urlstring(required)Starting URL to crawl
limitnumber10Max pages to crawl (up to 100,000)
depthnumber100,000Max link depth from starting URL
formatsarray["html"]Output formats: html, markdown, json
renderbooleantruetrue = headless browser, false = fast HTML fetch
sourcestring"all"Page discovery: all, sitemaps, links
maxAgenumber86400Cache validity in seconds (max 604800)
modifiedSincenumber-Unix timestamp; only crawl pages modified after this time
Options Object
ParameterTypeDefaultDescription
includePatternsarray[]Wildcard patterns to include (* and **)
excludePatternsarray[]Wildcard patterns to exclude (higher priority)
includeSubdomainsbooleanfalseFollow links to subdomains
includeExternalLinksbooleanfalseFollow external links
Show full SKILL.md (312 more words)Show less
Advanced Parameters
ParameterTypeDescription
jsonOptionsobjectAI-powered structured extraction (prompt, response_format)
authenticateobjectHTTP basic auth (username, password)
setExtraHTTPHeadersobjectCustom headers for requests
rejectResourceTypesarraySkip: image, media, font, stylesheet
userAgentstringCustom user agent string
cookiesarrayCustom cookies for requests

Usage Examples

Crawl documentation site (most common)
/cf-crawl https://docs.example.com --limit 50

Crawls up to 50 pages, saves as markdown.

Crawl with filters
/cf-crawl https://docs.example.com --limit 100 --include "/guides/**,/api/**" --exclude "/changelog/**"
Incremental crawl (diff detection)
/cf-crawl https://docs.example.com --limit 50 --since 2026-03-10

Only crawls pages modified since the given date. Skipped pages appear with status=skipped in results. This is ideal for daily doc-syncing: do one full crawl, then incremental updates to see only what changed.

Fast crawl without JavaScript rendering
/cf-crawl https://docs.example.com --no-render --limit 200

Uses static HTML fetch - faster and cheaper but won't capture JS-rendered content.

Crawl and merge into single file
/cf-crawl https://docs.example.com --limit 50 --merge

Merges all pages into a single markdown file for easy context loading.

Argument Parsing

When invoked as /cf-crawl, parse the arguments as follows:

  • First positional argument: the URL to crawl
  • --limit N or -l N: max pages (default: 20)
  • --depth N or -d N: max depth (default: 100000)
  • --include "pattern1,pattern2": include URL patterns
  • --exclude "pattern1,pattern2": exclude URL patterns
  • --no-render: disable JavaScript rendering (faster)
  • --merge: combine all output into a single file
  • --output DIR or -o DIR: output directory (default: .crawl-output)
  • --source sitemaps|links|all: page discovery method (default: all)
  • --since DATE: only crawl pages modified since DATE (ISO date like 2026-03-10 or Unix timestamp). Converts to Unix timestamp for the modifiedSince API parameter

If no URL is provided, ask the user for the target URL.

Important Notes

  • The /crawl endpoint respects robots.txt directives including crawl-delay
  • Blocked URLs appear with "status": "disallowed" in results
  • Free plan: 10 minutes of browser time per day
  • Job results are available for 14 days after completion
  • Max job runtime: 7 days
  • Response page size limit: 10 MB per page
  • Use render: false for static sites to save browser time
  • Pattern wildcards: * matches any character except /, ** matches including /

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-tool/components/skills/utilities/cf-crawl of davila7/claude-code-templates.

Open the folder on GitHubat commit 46b4d8b

Compare with similar skills

Cf Crawl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cf Crawl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cf Crawl this skilldavila7/claude-code-templates32k—~2.6kAutomated safety check: NotesMIT
Firecrawl Knowledge Basefirecrawl/skills117—~612Automated safety check: PassISC
Firecrawl Knowledge Ingestfirecrawl/skills117—~565Automated safety check: PassISC
Kb Refreshtechwolf-ai/ai-first-toolkit132—~1.3kAutomated safety check: PassMIT
OpenkbVectifyAI/OpenKB4.8k1 repos~2kAutomated safety check: WarnApache-2.0
Tencent ima Knowledge Base Readerzj-unicom-ai/UniEmployee359—~1kAutomated safety check: PassMIT

Similar skills

  • Firecrawl Knowledge Base

    firecrawl/skills

    Build a knowledge base from web content with Firecrawl. An agent skill from firecrawl/skills.

    117 GitHub stars~612 tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser.

    117 GitHub stars~565 tokensUpdated today
    Sales & SupportAuto-check passed
  • Kb Refresh

    techwolf-ai/ai-first-toolkit

    Add new sources to your knowledge base or re-scrape existing ones to pick up changes.

    132 GitHub stars~1.3k tokensUpdated 9 days ago
    Knowledge ManagementAuto-check passed
  • Openkb

    VectifyAI/OpenKB

    A skill your agent uses when the user asks about content in their OpenKB knowledge base — research topics, concepts compiled from their documents, cross-document synthesis — or mentions openkb, an…

    4.8k GitHub starsUsed in 1 repo~2k tokens
    Knowledge ManagementAuto-check: warnings
  • Tencent ima Knowledge Base Reader

    zj-unicom-ai/UniEmployee

    Exports the article list and original article text from a Tencent ima knowledge base through a logged-in Chrome session, using browser automation.

    359 GitHub stars~1k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Webcrawler Deep Crawl

    browser-act/skills

    Deep-crawl any website from start URLs, return per-page LLM-ready text/markdown/HTML plus metadata (title, description, author, language, canonical URL, OG) and in-scope outbound links.

    6.1k GitHub stars~4.3k tokensUpdated 1 mo ago
    DatabasesAuto-check passed

More from davila7/claude-code-templates

All 478 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 9 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Works with

Questions about Cf Crawl

What does Cf Crawl do?

Crawl entire websites using Cloudflare Browser Rendering /crawl API. Cf Crawl is an agent skill from davila7/claude-code-templates. Crawl entire websites using Cloudflare Browser Rendering /crawl API.

When should I use Cf Crawl?

Cf Crawl fits situations like: tasks that involve Web scraping; tasks that involve Static sites and blogs; tasks that involve Knowledge bases.

How do I install Cf Crawl in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill cf-crawl -a claude-code`. Or copy the skill folder (cli-tool/components/skills/utilities/cf-crawl in davila7/claude-code-templates) into .claude/skills/cf-crawl in your project. Claude Code loads it when a task matches its description.

How do I install Cf Crawl in Codex?

Run `npx skills add davila7/claude-code-templates --skill cf-crawl -a codex`. Or copy the skill folder (cli-tool/components/skills/utilities/cf-crawl in davila7/claude-code-templates) into .agents/skills/cf-crawl in your project. Codex loads it when a task matches its description.

Can I use Cf Crawl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill cf-crawl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cf-crawl, .gemini/skills/cf-crawl, .github/skills/cf-crawl and .opencode/skills/cf-crawl in your project.

What does Cf Crawl need to run?

Going by SKILL.md and its folder, Cf Crawl needs the command-line tools its instructions call (curl, python3 and cursor) and credentials named CLOUDFLARE_API_TOKEN and API_TOKEN. Our summary lists: Python 3; A credential in CLOUDFLARE_API_TOKEN; A credential in API_TOKEN.

Does Cf Crawl access the network?

SKILL.md names 2 domains. In commands or code: api.cloudflare.com; the agent is likely to contact it when it follows the instructions. As links in the text: dash.cloudflare.com. This is read from the text; nothing was executed.

Is Cf Crawl safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cf Crawl use?

Cf Crawl is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cf Crawl use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cf Crawl?

Skills that share tags, products or a category with Cf Crawl: Firecrawl Knowledge Base (firecrawl/skills, 117 stars), Firecrawl Knowledge Ingest (firecrawl/skills, 117 stars), Kb Refresh (techwolf-ai/ai-first-toolkit, 132 stars) and Openkb (VectifyAI/OpenKB, 4.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cf Crawl?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,483 GitHub stars. The repository holds 478 skills in this directory. The repository was last updated on October 9, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.