Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server.

MITAuto-check passedData & Analytics

Install Crawlberg

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill crawlberg -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins crawlberg --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/crawlberg .claude/skills/crawlberg && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawlberg
GitHub stars
1.2k
Token cost
~3.4k tokens
SKILL.md length
1,046 words
Files
1
Skills in repo
686
Repo updated
First seen
Licence
MIT

At a glance

Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server.

  • Works in 3 steps: Fetches statically via reqwest. → Detects WAF blocks (8 vendors) and… → Re-fetches through headless Chrome with…
  • The user wants to fetch a page
  • SKILL.md covers Installation, Command map, Scrape a single page and Crawl a site, plus 7 more sections
  • Calls brew, npx and uvx

What it does

Crawlberg is an agent skill from hashgraph-online/awesome-codex-plugins. Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Use when the user wants to fetch a page, follow links across a domain, enumerate URLs, or drive a real browser. Covers installation, the subcommands (scrape, crawl, map, interact, batch-scrape, batch-crawl, download, citations, version, mcp, serve), output formats (JSON + Markdown), browser fallback, and when to prefer the MCP server over shelling out.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping and MCP servers. It works with Model Context Protocol. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is MIT.

When your agent uses it

  • The user wants to fetch a page
  • Follow links across a domain
  • Drive a real browser

Example prompts

  • “/crawlberg”

Requirements

  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Fetches statically via reqwest.
  2. Detects WAF blocks (8 vendors) and JS-only shells.
  3. Re-fetches through headless Chrome with a real fingerprint when needed.

What it can do on your machine

Read from SKILL.md and the folder at commit 78497e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • brew
    • npx
    • uvx
    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and uvx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawlberg loads about 3.4k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 78497e5, republished under its MIT licence (© hashgraph-online). 1,046 words, ~3,357 tokens.

Download SKILL.mdSave it as .claude/skills/crawlberg/SKILL.md (or your agent's skills folder).
name
crawlberg
description
Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Use when the user wants to fetch a page, follow links across a domain, enumerate URLs, or drive a real browser. Covers installation, the subcommands (scrape, crawl, map, interact, batch-scrape, batch-crawl, download, citations, version, mcp, serve), output formats (JSON + Markdown), browser fallback, and when to prefer the MCP server over shelling out.
license
MIT
metadata.author
xberg-io
metadata.version
0.1.0
metadata.repository
https://github.com/xberg-io/crawlberg
<!--
AI-RULEZ :: GENERATED FILE — DO NOT EDIT
Content-Hash: blake3:2400d592929aaf450317f228f695638b76d0f7987f5996106372e06b9abce70d
Source-Hash: blake3:b5689383a3914da8e3cc6ad3614356ae7f98c90cf99cb1b4b47be38456ff7a7c
Schema-Version: v1
-->

Crawlberg

Crawlberg is a Rust-native web crawler and scraper. It fetches static HTML with reqwest, falls back to headless Chrome when a page needs JS or trips a WAF, and converts every result to clean Markdown via the built-in HTML→Markdown engine.

Use this skill when the user wants to:

  • Scrape a single URL to Markdown plus structured metadata.
  • Crawl a site following links bounded by depth, page count, and concurrency.
  • Enumerate URLs from sitemaps without paying for rendering.
  • Drive a real browser (click, type, scroll) and capture the resulting DOM.
  • Run the same operations from another agent harness via MCP tools.

Installation

The plugin shells out to a crawlberg binary on PATH. Install one of:

bash
brew install xberg-io/tap/crawlberg
# or run without a persistent install (the CLI proxy package self-installs the binary):
npx @xberg-io/crawlberg-cli --help
uvx --from crawlberg-cli crawlberg --help
# or build from source:
cargo install crawlberg-cli --features all

The serve and mcp subcommands are gated behind non-default cargo features (api and mcp). The Homebrew tap is built with all features, so both subcommands work out of the box. A from-source build must pass --features mcp (and --features api for serve), or --features all, to include them.

Verify:

bash
crawlberg --version

Headless fallback needs Chrome/Chromium reachable locally (chromiumoxide launches it on demand). Skip the install if you only plan to use --browser-mode never.

Command map

text
crawlberg scrape <url>          # single page → JSON or Markdown
crawlberg crawl <url...>        # follow links, BFS, depth-bounded
crawlberg map <url>             # enumerate URLs via sitemaps + link extraction
crawlberg interact <url>        # browser actions: click, type, scroll
crawlberg batch-scrape <url...> # scrape many URLs concurrently
crawlberg batch-crawl <url...>  # crawl many seed URLs concurrently
crawlberg download <url>        # download a document, report file metadata
crawlberg citations <input>     # markdown links → numbered citations (text or @file.md)
crawlberg version               # print the crawlberg version as JSON
crawlberg mcp                   # MCP server (stdio) — auto-registered (`mcp` feature)
crawlberg serve                 # REST API server (`api` feature)

crawl also handles batching implicitly: pass multiple seed URLs and it fans out via batch_crawl internally. The explicit batch-scrape and batch-crawl subcommands expose the same concurrency for many independent URLs.

Per-subcommand flags:

SubcommandPositionalKey flags
scrape<url>--proxy, --user-agent (plus shared flags below)
crawl<url...>--depth/-d (2), --max-pages/-n, --concurrent/-c (10), --rate-limit (200), --stay-on-domain, --proxy, --user-agent
map<url>--limit, --search
interact<url>--actions <json> (required)
batch-scrape<url...>--concurrent/-c (10), --proxy, --user-agent
batch-crawl<url...>--depth/-d (2), --max-pages/-n, --concurrent/-c (10), --rate-limit (200), --stay-on-domain, --proxy, --user-agent
download<url>--max-size
citations<input>none (input is markdown text or @file.md)
version—none
serve—--host (0.0.0.0), --port (3000)
mcp—none (stdio transport)
Shared flags
FlagDefaultNotes
--formatjsonjson or markdown.
--timeout30000Request timeout in milliseconds.
--browser-modeautoauto, always, or never.
--browser-endpoint—Optional CDP ws:// or wss:// URL.
--respect-robots-txtoffPass to obey robots.txt.
--config <json>—Inline JSON or @file.json to override defaults.

The --config flag accepts the full CrawlConfig schema. Anything you set explicitly on the CLI overrides the corresponding JSON field.

These shared flags apply to the crawl/scrape-family subcommands. --format, --browser-mode, --browser-endpoint, and --config cover scrape, crawl, map, interact, batch-scrape, and batch-crawl; download takes --timeout, --browser-mode, --browser-endpoint, --max-size, and --config (no --format). --respect-robots-txt applies to scrape, crawl, map, batch-scrape, and batch-crawl. citations and version take no shared flags.

Scrape a single page

bash
crawlberg scrape https://example.com --format markdown

JSON output (default) carries the rendered Markdown, page metadata (PageMetadata), links by category, images, feeds, JSON-LD blocks, and HTTP response metadata. Use Markdown output when piping into a file the user will read.

See the scraping-html-to-markdown skill for the full flag surface.

Crawl a site

bash
crawlberg crawl https://example.com \
  --depth 3 --max-pages 200 --concurrent 8 --rate-limit 250 \
  --stay-on-domain --respect-robots-txt --format markdown

Crawling is BFS by default, bounded by --depth, --max-pages, and --concurrent. Per-domain politeness is enforced by --rate-limit (milliseconds between requests to the same origin).

See the crawling-a-site skill for the recommended defaults and the full flag surface.

Map URLs

bash
crawlberg map https://example.com --limit 500 --search docs --format markdown

map reads sitemap.xml (and nested sitemaps), then falls back to link extraction from the seed page. It does not render pages — use it to plan a crawl or to feed URLs into another tool.

Browser interaction

bash
crawlberg interact https://example.com \
  --actions '[{"type":"click","selector":"#load-more"},
              {"type":"wait","milliseconds":500},
              {"type":"scrape"}]'

Action types are click, type, press, scroll, wait, screenshot, executeJs, and scrape (to wait for an element, use wait with a selector field). The result wraps the final HTML under interaction.final_html. See the automating-the-browser skill for the full action schema and limits.

Show full SKILL.md (475 more words)Show less

MCP server

When this plugin is installed in a Claude Code / Codex / Cursor / Gemini / opencode harness, the MCP server is auto-registered:

text
crawlberg mcp

mcp is a stdio-transport server and takes no arguments. It requires a binary built with the mcp feature (see Installation).

The server registers nine tools (the same set is served over the Streamable HTTP transport when running crawlberg serve):

ToolPurposeParameters
scrapeScrape one URL to Markdown or JSON (content, metadata, links).url (required), format (markdown|json), use_browser (bool — force browser)
crawlFollow links from a URL, bounded by depth/page count.url (required), max_depth, max_pages, format, stay_on_domain
mapDiscover all URLs via links and sitemaps.url (required), limit, search, respect_robots_txt
batch_scrapeScrape multiple URLs concurrently.urls (required array), format, concurrency
batch_crawlCrawl multiple seed URLs concurrently.urls (required array), max_depth, max_pages, format, stay_on_domain, concurrency
downloadDownload a document and return file metadata.url (required), max_size
interactExecute browser actions on a page (mutating/destructive).url (required), actions (required array of action objects)
generate_citationsRewrite markdown links as numbered citations + reference list.markdown (required)
get_versionReturn the crawlberg library version.none

Prefer MCP tools over shelling out when both are available:

  • Typed schemas surface argument errors before the call.
  • Results stream back as structured tool output instead of stdout text.
  • No --format juggling — the harness pulls whatever shape it needs.

Fall back to the CLI when you need to script a pipeline, capture stderr, or chain with shell tools.

Headless fallback

In --browser-mode auto (default), the engine:

  1. Fetches statically via reqwest.
  2. Detects WAF blocks (8 vendors) and JS-only shells.
  3. Re-fetches through headless Chrome with a real fingerprint when needed.

Force the browser path with --browser-mode always when you already know the page needs JS. Use --browser-mode never for hot loops where the cost of a stray Chrome launch is unacceptable.

Point --browser-endpoint ws://host:9222/devtools/browser/<id> at an already-running Chrome to skip the local launch.

See the headless-fallback skill for symptoms, costs, and external-CDP patterns.

Output formats

ModeUse when
jsonDownstream consumer needs metadata, links, images, etc.
markdownHuman reader or LLM-context payload.

Markdown output skips metadata. If you need both, run with --format json and read result.markdown.content.

Robots, rate limits, ethics

  • --respect-robots-txt is off by default; pass it for any crawl on a host you do not own.
  • The default --rate-limit 200 already produces a polite cadence; raise it for shared hosts.
  • Identify the crawler honestly via --user-agent. Do not impersonate a browser unless the operator has approved it.

Cross-references

  • skills/crawling-a-site/SKILL.md — multi-page crawl with depth, page caps, concurrency, rate limits, and domain scoping.
  • skills/scraping-html-to-markdown/SKILL.md — single-page rendering, the Markdown output shape, and common pitfalls.
  • skills/mapping-urls/SKILL.md — map: sitemap + link URL discovery, filtering, and seeding a crawl.
  • skills/automating-the-browser/SKILL.md — interact: the full scripted action schema, limits, and result shape.
  • skills/serving-the-api/SKILL.md — serve: the Firecrawl-v1-compatible REST API server and its endpoints.
  • skills/headless-fallback/SKILL.md — when and how to force the browser backend.

© hashgraph-online, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/crawlberg of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 78497e5

Compare with similar skills

Crawlberg next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawlberg compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawlberg this skillhashgraph-online/awesome-codex-plugins1.2k—~3.4kAutomated safety check: PassMIT
Querying Indonesian Gov Datasuryast/indonesia-gov-apis172—~997Automated safety check: PassMIT
Scraplingforyourhealth111-pixel/Vibe-Skills3.6k—~1.1kAutomated safety check: PassApache-2.0
Firecrawl MCPLeoYeAI/openclaw-master-skills2.2k—~7.4kAutomated safety check: PassMIT
Skill Seekers Builderyusufkaraaslan/Skill_Seekers15k—~760Automated safety check: PassMIT
Sandbaseiflytek/skillhub5.2k2 repos~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Querying Indonesian Gov Data

    suryast/indonesia-gov-apis

    Query 57 Indonesian government APIs and data sources — BPJPH halal certification, BPOM food safety, OJK financial legality, BPS statistics, BMKG weather/earthquakes, Bank Indonesia exchange rates…

    172 GitHub stars~997 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Scrapling

    foryourhealth111-pixel/Vibe-Skills

    CLI-first web scraping & content extraction with optional MCP server.

    3.6k GitHub stars~1.1k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Firecrawl MCP

    LeoYeAI/openclaw-master-skills

    Auto-generated skill for firecrawl-mcp tools via OneKey Gateway.

    2.2k GitHub stars~7.4k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Skill Seekers Builder

    yusufkaraaslan/Skill_Seekers

    Detects the type of a knowledge source and uses the Skill Seekers MCP tools to turn docs, repos, PDFs or videos into packaged AI skills.

    15k GitHub stars~760 tokensUpdated 8 days ago
    Agent WorkflowsAuto-check passed
  • Sandbase

    iflytek/skillhub

    Access 2,000+ AI models and API tools through one MCP interface for inference, media generation, search, scraping, embeddings, social data, and structured retrieval.

    5.2k GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • X Twitter Scraper

    Xquik-dev/x-twitter-scraper

    Use Xquik to fetch X (Twitter) data or act through a connected account: search, profiles, followers, replies, threads, timelines, media downloads, bulk exports, trends, monitors, signed webhooks…

    209 GitHub starsUsed in 1 repo~2.6k tokens
    Backend & APIsAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 686 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.2k GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.2k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.2k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.2k GitHub stars~618 tokensUpdated today
    Auto-check passed
  • Manuscript Engagement Analytics

    hashgraph-online/awesome-codex-plugins

    Analyze nonfiction manuscripts for reader engagement signals, including heading-level word counts, slow starts, long slogs, weak takeaway titles, value pacing, beta-reader comment dropoff, and…

    1.2k GitHub stars~875 tokensUpdated today
    Auto-check passed

Questions about Crawlberg

What does Crawlberg do?

Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Crawlberg is an agent skill from hashgraph-online/awesome-codex-plugins. Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server.

When should I use Crawlberg?

Crawlberg fits situations like: the user wants to fetch a page; follow links across a domain; drive a real browser.

How do I install Crawlberg in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill crawlberg -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/crawlberg in hashgraph-online/awesome-codex-plugins) into .claude/skills/crawlberg in your project. Claude Code loads it when a task matches its description.

How do I install Crawlberg in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill crawlberg -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/crawlberg in hashgraph-online/awesome-codex-plugins) into .agents/skills/crawlberg in your project. Codex loads it when a task matches its description.

Can I use Crawlberg in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill crawlberg -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawlberg, .gemini/skills/crawlberg, .github/skills/crawlberg and .opencode/skills/crawlberg in your project.

What does Crawlberg need to run?

Going by SKILL.md and its folder, Crawlberg needs the command-line tools its instructions call (brew, npx, uvx and cargo). Our summary lists: Node.js.

Does Crawlberg access the network?

SKILL.md contains no URLs. Its commands use npx and uvx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Crawlberg safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Crawlberg use?

Crawlberg is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawlberg use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Crawlberg?

Skills that share tags, products or a category with Crawlberg: Querying Indonesian Gov Data (suryast/indonesia-gov-apis, 172 stars), Scrapling (foryourhealth111-pixel/Vibe-Skills, 3.6k stars), Firecrawl MCP (LeoYeAI/openclaw-master-skills, 2.2k stars) and Skill Seekers Builder (yusufkaraaslan/Skill_Seekers, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawlberg?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,242 GitHub stars. The repository holds 686 skills in this directory. The repository was last updated on October 8, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.