Monitor With Haoleme
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
$ npx skills add smallnest/goclaw --skill crawl4ai -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install smallnest/goclaw crawl4ai --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .claude/skills/crawl4ai && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .claude/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add smallnest/goclaw --skill crawl4ai -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install smallnest/goclaw crawl4ai --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .agents/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .agents/skills/crawl4ai && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .agents/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add smallnest/goclaw --skill crawl4ai -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install smallnest/goclaw crawl4ai --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .cursor/skills/crawl4ai && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .cursor/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/smallnest/goclaw.git --path internal/builtin_skills/crawl4ai-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add smallnest/goclaw --skill crawl4ai -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install smallnest/goclaw crawl4ai --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .gemini/skills/crawl4ai && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .gemini/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install smallnest/goclaw crawl4aiInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add smallnest/goclaw --skill crawl4ai -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .github/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .github/skills/crawl4ai && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .github/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add smallnest/goclaw --skill crawl4ai -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install smallnest/goclaw crawl4ai --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smallnest/goclaw.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/internal/builtin_skills/crawl4ai-skill .opencode/skills/crawl4ai && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "crawl4ai" agent skill from https://github.com/smallnest/goclaw/tree/master/internal/builtin_skills/crawl4ai-skill into .opencode/skills/crawl4ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawl4ai", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
crawl4aiScrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
The skill documents two interfaces to Crawl4AI: the crwl command line, recommended for quick scriptable jobs, and the Python SDK for full programmatic control, with separate CLI and SDK guides. Both share the same configuration layers: browser settings such as headless mode, viewport, user agent and proxy; crawler settings such as page timeout, wait conditions, cache mode, JavaScript to run and CSS selector focus; extraction from a schema; and content filters.
Every crawl returns markdown, raw HTML, discovered links, media and any extracted content. Clean markdown generation is the primary use case, with relevance filters such as BM25. Bundled Python scripts cover a basic crawler, batch crawling, an extraction pipeline and a Google search helper, and a tests folder checks crawling, extraction and markdown generation. Installation is pip install crawl4ai followed by crawl4ai-setup.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e05c79d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
shop.comsite1.comsite2.comsite3.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Crawl4AI Web Scraping loads about 2.5k tokens when it runs, and up to ~63k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 439 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from smallnest/goclaw at commit e05c79d, republished under its MIT licence (© smallnest). 439 words, ~2,466 tokens.
.claude/skills/crawl4ai/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.Crawl4AI provides comprehensive web crawling and data extraction capabilities. This skill supports both CLI (recommended for quick tasks) and Python SDK (for programmatic control).
Choose your interface:
crwl) - Quick, scriptable commands: CLI Guidepip install crawl4ai
crawl4ai-setup
# Verify installation
crawl4ai-doctor# Basic crawling - returns markdown
crwl https://example.com
# Get markdown output
crwl https://example.com -o markdown
# JSON output with cache bypass
crwl https://example.com -o json -v --bypass-cache
# See more examples
crwl --exampleimport asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://example.com")
print(result.markdown[:500])
asyncio.run(main())For SDK configuration details: SDK Guide - Configuration (lines 61-150)
Both CLI and SDK use the same underlying configuration:
| Concept | CLI | SDK |
|---|---|---|
| Browser settings | -B browser.yml or -b "param=value" | BrowserConfig(...) |
| Crawl settings | -C crawler.yml or -c "param=value" | CrawlerRunConfig(...) |
| Extraction | -e extract.yml -s schema.json | extraction_strategy=... |
| Content filter | -f filter.yml | markdown_generator=... |
Browser Configuration:
headless: Run with/without GUIviewport_width/height: Browser dimensionsuser_agent: Custom user agentproxy_config: Proxy settingsCrawler Configuration:
page_timeout: Max page load time (ms)wait_for: CSS selector or JS condition to wait forcache_mode: bypass, enabled, disabledjs_code: JavaScript to executecss_selector: Focus on specific elementFor complete parameters: CLI Config | SDK Config
Every crawl returns:
Crawl4AI excels at generating clean, well-formatted markdown:
# Basic markdown
crwl https://docs.example.com -o markdown
# Filtered markdown (removes noise)
crwl https://docs.example.com -o markdown-fit
# With content filter
crwl https://docs.example.com -f filter_bm25.yml -o markdown-fitFilter configuration:
# filter_bm25.yml (relevance-based)
type: "bm25"
query: "machine learning tutorials"
threshold: 1.0from crawl4ai.content_filter_strategy import BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
bm25_filter = BM25ContentFilter(user_query="machine learning", bm25_threshold=1.0)
md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)
config = CrawlerRunConfig(markdown_generator=md_generator)
result = await crawler.arun(url, config=config)
print(result.markdown.fit_markdown) # Filtered
print(result.markdown.raw_markdown) # OriginalFor content filters: Content Processing (lines 2481-3101)
No LLM required - fast, deterministic, cost-free.
CLI:
# Generate schema once (uses LLM)
python scripts/extraction_pipeline.py --generate-schema https://shop.com "extract products"
# Use schema for extraction (no LLM)
crwl https://shop.com -e extract_css.yml -s product_schema.json -o jsonSchema format:
{
"name": "products",
"baseSelector": ".product-card",
"fields": [
{"name": "title", "selector": "h2", "type": "text"},
{"name": "price", "selector": ".price", "type": "text"},
{"name": "link", "selector": "a", "type": "attribute", "attribute": "href"}
]
}For complex or irregular content:
CLI:
# extract_llm.yml
type: "llm"
provider: "openai/gpt-4o-mini"
instruction: "Extract product names and prices"
api_token: "your-token"crwl https://shop.com -e extract_llm.yml -o jsonFor extraction details: Extraction Strategies (lines 4522-5429)
CLI:
crwl https://example.com -c "wait_for=css:.ajax-content,scan_full_page=true,page_timeout=60000"Crawler config:
# crawler.yml
wait_for: "css:.ajax-content"
scan_full_page: true
page_timeout: 60000
delay_before_return_html: 2.0CLI (sequential):
for url in url1 url2 url3; do crwl "$url" -o markdown; donePython SDK (concurrent):
urls = ["https://site1.com", "https://site2.com", "https://site3.com"]
results = await crawler.arun_many(urls, config=config)For batch processing: arun_many() Reference (lines 1057-1224)
CLI:
# login_crawler.yml
session_id: "user_session"
js_code: |
document.querySelector('#username').value = 'user';
document.querySelector('#password').value = 'pass';
document.querySelector('#submit').click();
wait_for: "css:.dashboard"# Login
crwl https://site.com/login -C login_crawler.yml
# Access protected content (session reused)
crwl https://site.com/protected -c "session_id=user_session"For session management: Advanced Features (lines 5429-5940)
CLI:
# browser.yml
headless: true
proxy_config:
server: "http://proxy:8080"
username: "user"
password: "pass"
user_agent_mode: "random"crwl https://example.com -B browser.yml# Search Google and get results as JSON
python scripts/google_search.py "your search query" 20
# Example
python scripts/google_search.py "2026年Go语言展望" 20The script extracts:
Output is saved to google_search_results.json and printed to stdout.
crwl https://docs.example.com -o markdown > docs.md# Generate schema once
python scripts/extraction_pipeline.py --generate-schema https://shop.com "extract products"
# Monitor (no LLM costs)
crwl https://shop.com -e extract_css.yml -s schema.json -o json# Multiple sources with filtering
for url in news1.com news2.com news3.com; do
crwl "https://$url" -f filter_bm25.yml -o markdown-fit
done# First view content
crwl https://example.com -o markdown
# Then ask questions
crwl https://example.com -q "What are the main conclusions?"
crwl https://example.com -q "Summarize the key points"| Document | Purpose |
|---|---|
| CLI Guide | Command-line interface reference |
| SDK Guide | Python SDK quick reference |
| Complete SDK Reference | Full API documentation (5900+ lines) |
--bypass-cache only when neededcrwl https://example.com -c "wait_for=css:.dynamic-content,page_timeout=60000"crwl https://example.com -B browser.yml# browser.yml
headless: false
viewport_width: 1920
viewport_height: 1080
user_agent: "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"# Debug: see full output
crwl https://example.com -o all -v
# Try different wait strategy
crwl https://example.com -c "wait_for=js:document.querySelector('.content')!==null"# Verify session
crwl https://site.com -c "session_id=test" -o all | grep -i sessionFor comprehensive API documentation, see Complete SDK Reference.
© smallnest, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 15 other files (scripts, references) in internal/builtin_skills/crawl4ai-skill of smallnest/goclaw.
Open the folder on GitHubat commit e05c79d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in smallnest/goclaw, which our catalogue first saw on October 7, 2026.
Crawl4AI Web Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Crawl4AI Web Scraping this skillsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Authoritative Data Harvesteryushui2022/MathModel-Skill | 452 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper | 1.5k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Axyusukebe/ax | 719 | 1 repos | ~918 | Automated safety check: Pass | MIT | |
| Google Maps ScraperMahanaicoach/google-maps-scraper-kit | 1.3k | — | ~2.8k | Automated safety check: Pass | MIT |
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
eatmoreduck/boss-zhipin-scraper
Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.
yusukebe/ax
Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.
Mahanaicoach/google-maps-scraper-kit
Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
smallnest/goclaw
Reference for OpenClaw's discord tool: sending messages, reactions, stickers, emoji uploads, polls, threads and moderation in Discord channels and DMs.
smallnest/goclaw
Looks up action manuals with verified selectors for multi-step website tasks, then drives the browser with the actionbook CLI in a dedicated or an existing Chrome session.
smallnest/goclaw
Controls an Android device over ADB in a loop of screenshot, vision analysis, tap and verification, so the agent acts on real on-screen coordinates.
smallnest/goclaw
Upload images to Feishu/Lark. An agent skill from smallnest/goclaw.
smallnest/goclaw
Helps you find, install, update and remove skills in the goclaw agent framework, using the goclaw skills commands and outside skill sources.
Works with
Categories
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM. The skill documents two interfaces to Crawl4AI: the crwl command line, recommended for quick scriptable jobs, and the Python SDK for full programmatic control, with separate CLI and SDK guides. Both share the same configuration layers: browser settings such as headless mode, viewport, user agent and proxy; crawler settings such as page timeout, wait conditions, cache mode, JavaScript to run and CSS selector focus; extraction from a schema; and content filters.
Crawl4AI Web Scraping fits situations like: scraping a website into clean markdown; extracting structured data from JavaScript-heavy pages; crawling many URLs in a batch; building an automated web data pipeline.
Run `npx skills add smallnest/goclaw --skill crawl4ai -a claude-code`. Or copy the skill folder (internal/builtin_skills/crawl4ai-skill in smallnest/goclaw) into .claude/skills/crawl4ai in your project. Claude Code loads it when a task matches its description.
Run `npx skills add smallnest/goclaw --skill crawl4ai -a codex`. Or copy the skill folder (internal/builtin_skills/crawl4ai-skill in smallnest/goclaw) into .agents/skills/crawl4ai in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add smallnest/goclaw --skill crawl4ai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawl4ai, .gemini/skills/crawl4ai, .github/skills/crawl4ai and .opencode/skills/crawl4ai in your project.
Going by SKILL.md and its folder, Crawl4AI Web Scraping needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python with the crawl4ai package, installed with pip and crawl4ai-setup.
SKILL.md names 4 domains. In commands or code: shop.com, site1.com, site2.com and site3.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Crawl4AI Web Scraping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 61k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Crawl4AI Web Scraping: Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars) and Ax (yusukebe/ax, 719 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
smallnest (a GitHub user) maintains it in smallnest/goclaw, which has 598 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on March 27, 2026.
Source: smallnest/goclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.