Apify Actor Development
apify/agent-skills
Creates, changes, debugs and deploys Apify Actors, including their input and output schemas, using the Apify CLI.
Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/crawlee .claude/skills/crawlee && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .claude/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawleeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/crawlee .agents/skills/crawlee && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .agents/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/crawlee .cursor/skills/crawlee && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .cursor/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/crawlee--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/crawlee .gemini/skills/crawlee && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .gemini/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills crawleeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/crawlee .github/skills/crawlee && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .github/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/crawlee .opencode/skills/crawlee && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "crawlee" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/crawlee into .opencode/skills/crawlee/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crawlee", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
crawleeExpert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).
Crawlee is an agent skill from LeoYeAI/openclaw-master-skills. Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python). Use this skill whenever the user wants to: scrape a website, build a web crawler, extract data from web pages, automate browser navigation, handle anti-bot blocking, manage proxies or sessions for scraping, use Playwright/Puppeteer/Cheerio/BeautifulSoup for web data extraction, crawl sitemaps, download files from URLs, or deploy a scraper to the cloud. Trigger even for loosely related phrases like "get data from…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `_meta.json`, `references/js-api.md` and `references/python-api.md`).
It sits in Data & Analytics, covering Web scraping. It works with Python, JavaScript, TypeScript and Playwright. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmpipnpxplaywrightFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
proxy1.comproxy2.comAlso links to:
crawlee.devgithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Crawlee loads about 4.7k tokens when it runs, and up to ~8.2k if it reads all its reference files. Until then it costs about 198 tokens; SKILL.md has 478 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 478 words, ~4,725 tokens.
.claude/skills/crawlee/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Crawlee is a production-grade web scraping and browser automation library for JavaScript/TypeScript (Node.js 16+) and Python (3.10+). It handles anti-blocking, proxies, session management, storage, and concurrency out of the box.
Docs: https://crawlee.dev/js/docs | https://crawlee.dev/python/docs
GitHub: https://github.com/apify/crawlee
| Crawler | When to Use | JS Required |
|---|---|---|
CheerioCrawler | Fast HTML parsing, no JS rendering needed | ❌ |
HttpCrawler | Raw HTTP responses, custom parsing | ❌ |
JSDOMCrawler | DOM manipulation without full browser | ❌ |
PlaywrightCrawler | Modern headless browser (Chromium/Firefox/WebKit) | ✅ |
PuppeteerCrawler | Chromium/Chrome headless automation | ✅ |
AdaptivePlaywrightCrawler | Auto-detects if JS rendering is needed | Auto |
BasicCrawler | Custom HTTP logic from scratch | ❌ |
Rule of thumb: Start with CheerioCrawler. Upgrade to PlaywrightCrawler only when JS rendering is required.
| Crawler | When to Use |
|---|---|
BeautifulSoupCrawler | HTML parsing with BeautifulSoup (fast, no JS) |
ParselCrawler | CSS/XPath selectors, Scrapy-style (fast, no JS) |
PlaywrightCrawler | Full browser automation (Chromium/Firefox/WebKit) |
AdaptivePlaywrightCrawler | Auto HTTP vs browser decision |
# Recommended: use the CLI
npx crawlee create my-crawler
cd my-crawler && npm install
# Or manually:
npm install crawlee
# For Playwright:
npm install crawlee playwright
npx playwright install
# For Puppeteer:
npm install crawlee puppeteerAdd to package.json:
{ "type": "module" }pip install crawlee
# With BeautifulSoup:
pip install 'crawlee[beautifulsoup]'
# With Playwright:
pip install 'crawlee[playwright]'
playwright installRequest objects in a RequestQueuerequestHandler function (JS) / decorated handler (Python)Request — A single URL + metadata to crawlRequestQueue — Dynamic, deduplicated queue of URLsDataset — Append-only structured result storage (like a table)KeyValueStore — Blob storage for screenshots, PDFs, stateProxyConfiguration — Manages proxy rotationSessionPool — Manages browser sessions + cookiesimport { CheerioCrawler, Dataset } from 'crawlee';
const crawler = new CheerioCrawler({
async requestHandler({ $, request, enqueueLinks, log }) {
const title = $('title').text();
log.info(`Title of ${request.loadedUrl}: ${title}`);
await Dataset.pushData({ url: request.loadedUrl, title });
// Enqueue all links found on this page
await enqueueLinks();
},
maxRequestsPerCrawl: 100, // Safety limit
});
await crawler.run(['https://example.com']);import { PlaywrightCrawler, Dataset } from 'crawlee';
const crawler = new PlaywrightCrawler({
// headless: false, // Uncomment to see the browser
async requestHandler({ page, request, enqueueLinks, log }) {
const title = await page.title();
log.info(`${request.loadedUrl}: ${title}`);
await Dataset.pushData({ url: request.loadedUrl, title });
await enqueueLinks();
},
});
await crawler.run(['https://example.com']);import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main() -> None:
crawler = BeautifulSoupCrawler(max_requests_per_crawl=50)
@crawler.router.default_handler
async def handler(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.string if context.soup.title else None
context.log.info(f'Processing {context.request.url}: {title}')
await context.push_data({'url': context.request.url, 'title': title})
await context.enqueue_links()
await crawler.run(['https://example.com'])
if __name__ == '__main__':
asyncio.run(main())import asyncio
from crawlee.crawlers import PlaywrightCrawler, PlaywrightCrawlingContext
async def main() -> None:
crawler = PlaywrightCrawler(headless=True, browser_type='chromium')
@crawler.router.default_handler
async def handler(context: PlaywrightCrawlingContext) -> None:
title = await context.page.title()
await context.push_data({'url': context.request.url, 'title': title})
await context.enqueue_links()
await crawler.run(['https://example.com'])
if __name__ == '__main__':
asyncio.run(main())Use labels + router to handle different kinds of pages (list pages, detail pages, etc.).
import { PlaywrightCrawler, Dataset } from 'crawlee';
import { router } from './routes.js';
const crawler = new PlaywrightCrawler({ requestHandler: router });
await crawler.run([{ url: 'https://shop.example.com', label: 'START' }]);// routes.js
import { createPlaywrightRouter } from 'crawlee';
export const router = createPlaywrightRouter();
router.addHandler('START', async ({ page, enqueueLinks }) => {
await enqueueLinks({ selector: 'a.category', label: 'CATEGORY' });
});
router.addHandler('CATEGORY', async ({ page, enqueueLinks }) => {
await enqueueLinks({ selector: 'a.product', label: 'DETAIL' });
// Enqueue next page
const next = await page.$('a.next-page');
if (next) await enqueueLinks({ selector: 'a.next-page', label: 'CATEGORY' });
});
router.addDefaultHandler(async ({ page, request, pushData }) => {
// DETAIL pages
const title = await page.title();
const price = await page.$eval('.price', el => el.textContent);
await pushData({ url: request.url, title, price });
});from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
crawler = BeautifulSoupCrawler()
@crawler.router.handler('CATEGORY')
async def category_handler(context: BeautifulSoupCrawlingContext) -> None:
await context.enqueue_links(selector='a.product', label='DETAIL')
@crawler.router.default_handler
async def detail_handler(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.string
await context.push_data({'url': context.request.url, 'title': title})enqueueLinks()// Enqueue all links on page
await enqueueLinks();
// Filter by glob pattern
await enqueueLinks({ globs: ['https://example.com/products/**'] });
// Filter by regex
await enqueueLinks({ regexps: [/\/product\/\d+/] });
// Enqueue only specific selector
await enqueueLinks({ selector: 'a.pagination', label: 'LIST' });
// Enqueue with custom label and transformations
await enqueueLinks({
selector: 'a.item',
label: 'DETAIL',
transformRequestFunction: (req) => {
req.userData.scrapedAt = new Date().toISOString();
return req;
},
});await context.enqueue_links()
await context.enqueue_links(selector='a.product', label='DETAIL')
await context.enqueue_links(include=[re.compile(r'/products/\d+')])// JS — Write
await Dataset.pushData({ url, title, price });
await Dataset.pushData([item1, item2, item3]); // batch write
// JS — Read / Export
const dataset = await Dataset.open();
await dataset.exportToCSV('results'); // saves to KV store
await dataset.exportToJSON('results');
for await (const item of dataset) { console.log(item); }# Python — Write
await context.push_data({'url': url, 'title': title})
# Python — Read / Export
from crawlee.storages import Dataset
dataset = await Dataset.open()
await dataset.export_to(key='results', content_type='csv')Data is saved to ./storage/datasets/default/*.json by default.
// JS
await KeyValueStore.setValue('OUTPUT', { results: [...] });
const value = await KeyValueStore.getValue('OUTPUT');
// Save a screenshot
const store = await KeyValueStore.open();
await store.setValue('screenshot', await page.screenshot(), { contentType: 'image/png' });# Python
from crawlee.storages import KeyValueStore
kvs = await KeyValueStore.open()
await kvs.set_value('result', {'data': 'value'})
value = await kvs.get_value('result')./storage/
datasets/default/ # Dataset rows as JSON files
key_value_stores/default/ # KV store entries
request_queues/default/ # Request queue stateOverride with env var: CRAWLEE_STORAGE_DIR=/path/to/storage
// JS — Basic proxy rotation
import { ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: [
'http://user:pass@proxy1.example.com:8000',
'http://user:pass@proxy2.example.com:8000',
],
});
const crawler = new CheerioCrawler({
proxyConfiguration,
useSessionPool: true,
persistCookiesPerSession: true,
async requestHandler({ proxyInfo, request }) {
console.log('Using proxy:', proxyInfo?.url);
},
});// JS — Tiered proxies (smart cost/reliability balancing)
const proxyConfiguration = new ProxyConfiguration({
tieredProxyUrls: [
[null], // Tier 0: no proxy (cheapest)
['http://cheap-datacenter-proxy'], // Tier 1: datacenter
['http://expensive-residential'], // Tier 2: residential (most reliable)
],
});
// Crawlee auto-escalates tiers when blocking is detected, then drops back when clear# Python
from crawlee.proxy_configuration import ProxyConfiguration
proxy_configuration = ProxyConfiguration(
proxy_urls=['http://proxy1.com/', 'http://proxy2.com/'],
)
crawler = BeautifulSoupCrawler(
proxy_configuration=proxy_configuration,
use_session_pool=True,
)Sessions tie together cookies, proxy IPs, and headers to simulate a consistent user identity.
// JS
const crawler = new CheerioCrawler({
useSessionPool: true, // Enable (default: true)
persistCookiesPerSession: true,
sessionPoolOptions: { maxPoolSize: 100 },
async requestHandler({ session, $ }) {
const title = $('title').text();
if (title === 'Access Denied') {
session?.retire(); // Mark this IP+cookie combo as blocked
} else if (title === 'Slow') {
session?.markBad(); // Penalize but don't retire
}
// session.markGood() is called automatically on success
},
});# Python
from crawlee.sessions import SessionPool
crawler = BeautifulSoupCrawler(
use_session_pool=True,
session_pool=SessionPool(max_pool_size=100),
)
@crawler.router.default_handler
async def handler(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.string if context.soup.title else ''
if title == 'Access Denied':
context.session.retire()// JS — Playwright with fingerprint rotation (built-in, zero config needed)
const crawler = new PlaywrightCrawler({
// Fingerprints automatically randomized by default in Playwright/Puppeteer crawlers
// headless: false, // Use headful for harder targets
async requestHandler({ page }) {
// Add realistic delays
await page.waitForTimeout(1000 + Math.random() * 2000);
},
});
// Use got-scraping for HTTP (built into CheerioCrawler/HttpCrawler)
// It automatically sets realistic headers and TLS fingerprintsAnti-blocking checklist:
CheerioCrawler — it uses got-scraping which mimics real browser HTTPuseSessionPool: true with a proxyConfigurationmaxRequestsPerMinute to avoid rate limitspersistCookiesPerSession: truesession.retire()// JS
const crawler = new CheerioCrawler({
maxConcurrency: 50, // Max parallel requests (default: 200)
minConcurrency: 1, // Don't set too high!
maxRequestsPerMinute: 120, // Rate limit
maxRequestsPerCrawl: 1000, // Total request cap (safety)
requestHandlerTimeoutSecs: 30,
});# Python
from crawlee import ConcurrencySettings
crawler = BeautifulSoupCrawler(
concurrency_settings=ConcurrencySettings(
max_concurrency=50,
max_tasks_per_minute=120,
),
max_requests_per_crawl=1000,
)Scaling notes:
minConcurrency high — it can crash under loadmaxRequestsPerMinute is smoother than raw concurrency throttling| Env Variable | Default | Purpose |
|---|---|---|
CRAWLEE_STORAGE_DIR | ./storage | Storage root directory |
CRAWLEE_DEFAULT_DATASET_ID | default | Override default dataset ID |
CRAWLEE_DEFAULT_KEY_VALUE_STORE_ID | default | Override default KVS ID |
CRAWLEE_DEFAULT_REQUEST_QUEUE_ID | default | Override default queue ID |
CRAWLEE_PURGE_ON_START | true | Clear storage before each run |
// JS — Programmatic configuration
import { Configuration } from 'crawlee';
const config = new Configuration({
storageDir: '/data/crawlee',
persistStateIntervalMillis: 30_000,
});
const crawler = new CheerioCrawler({ /* ... */ }, config);FROM apify/actor-node-playwright-chrome:20
COPY package*.json ./
RUN npm ci --only=prod
COPY . ./
CMD ["node", "src/main.js"]For Cheerio (smaller image):
FROM apify/actor-node:20// JS — Enqueue next page
router.addHandler('LIST', async ({ page, enqueueLinks }) => {
await enqueueLinks({ selector: '.product', label: 'DETAIL' });
const hasNext = await page.$('a.next');
if (hasNext) await enqueueLinks({ selector: 'a.next', label: 'LIST' });
});// JS — Save to KeyValueStore
const { body } = await sendRequest({ responseType: 'buffer' });
await KeyValueStore.setValue('file.pdf', body, { contentType: 'application/pdf' });// JS — Playwright
async requestHandler({ page, request }) {
const screenshot = await page.screenshot({ fullPage: true });
await KeyValueStore.setValue(
`screenshot-${Date.now()}`,
screenshot,
{ contentType: 'image/png' }
);
}// JS — useState()
async requestHandler({ useState }) {
const state = await useState({ count: 0 });
state.count++;
console.log('Total processed:', state.count);
}// JS
const crawler = new CheerioCrawler({
maxRequestRetries: 3, // Retry failed requests up to 3 times
failedRequestHandler: async ({ request, error }) => {
console.error(`Failed: ${request.url}`, error.message);
await Dataset.pushData({ url: request.url, error: error.message });
},
});# Python
crawler = BeautifulSoupCrawler(max_request_retries=3)
@crawler.failed_request_handler
async def on_failed(context: BasicCrawlingContext, error: Exception) -> None:
context.log.error(f'Failed {context.request.url}: {error}')import { CheerioCrawler } from 'crawlee';
import { Sitemap } from '@crawlee/utils';
const { urls } = await Sitemap.load('https://example.com/sitemap.xml');
const crawler = new CheerioCrawler({ /* ... */ });
await crawler.run(urls);import { CheerioCrawler } from 'crawlee';
import { createServer } from 'http';
const server = createServer(async (req, res) => {
const url = new URL(req.url, 'http://localhost').searchParams.get('url');
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 1,
async requestHandler({ $ }) {
res.end(JSON.stringify({ title: $('title').text() }));
},
});
await crawler.run([url]);
});
server.listen(3000);import { CheerioCrawler, CheerioCrawlingContext, Dataset } from 'crawlee';
interface Product {
url: string;
title: string;
price: number;
}
const crawler = new CheerioCrawler({
async requestHandler({ $, request }: CheerioCrawlingContext) {
const title = $('h1').text();
const price = parseFloat($('.price').text().replace('$', ''));
await Dataset.pushData<Product>({ url: request.url, title, price });
},
});import { Actor } from 'apify';
import { CheerioCrawler } from 'crawlee';
await Actor.init();
const input = await Actor.getInput();
const { startUrls } = input;
const crawler = new CheerioCrawler({
async requestHandler({ $, request }) {
await Actor.pushData({ url: request.url, title: $('title').text() });
},
});
await crawler.run(startUrls);
await Actor.exit();Deploy with: apify push
// Enable verbose logging
import { Log } from 'crawlee';
Log.setLevel(Log.LEVELS.DEBUG);
// Run headful (browser crawlers only)
const crawler = new PlaywrightCrawler({
headless: false,
// ...
});
// Limit requests while developing
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 10,
// ...
});For advanced topics, see:
references/js-api.md — Full JS API quick referencereferences/python-api.md — Full Python API quick referenceBoth language docs: https://crawlee.dev
© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/crawlee of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Crawlee next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Crawlee this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.7k | Automated safety check: Pass | MIT | |
| Apify Actor Developmentapify/agent-skills | 2.4k | — | ~2.9k | Automated safety check: Pass | None | |
| Brightdata SDK JSbrightdata/skills | 264 | — | ~3k | Automated safety check: Pass | MIT | |
| Anti Detect Browserantibrow/anti-detect-browser-skills | 17 | — | ~9.8k | Automated safety check: Warn | MIT | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Agent Browseroxylabs/agent-skills | 875 | — | ~3k | Automated safety check: Pass | MIT |
apify/agent-skills
Creates, changes, debugs and deploys Apify Actors, including their input and output schemas, using the Apify CLI.
brightdata/skills
Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).
antibrow/anti-detect-browser-skills
Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
oxylabs/agent-skills
Connects to Oxylabs remote agent browsers over the Chrome DevTools Protocol (CDP) with Playwright or Puppeteer.
brightdata/skills
Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Categories
Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python). Crawlee is an agent skill from LeoYeAI/openclaw-master-skills. Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).
Crawlee fits situations like: the user wants to: scrape a website; build a web crawler; extract data from web pages; automate browser navigation.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a claude-code`. Or copy the skill folder (skills/crawlee in LeoYeAI/openclaw-master-skills) into .claude/skills/crawlee in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a codex`. Or copy the skill folder (skills/crawlee in LeoYeAI/openclaw-master-skills) into .agents/skills/crawlee in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawlee, .gemini/skills/crawlee, .github/skills/crawlee and .opencode/skills/crawlee in your project.
Going by SKILL.md and its folder, Crawlee needs the command-line tools its instructions call (npm, pip, npx and playwright). Our summary lists: Python 3; Node.js.
SKILL.md names 4 domains. In commands or code: proxy1.com and proxy2.com; the agent is likely to contact these when it follows the instructions. As links in the text: crawlee.dev and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Crawlee is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Crawlee: Apify Actor Development (apify/agent-skills, 2.4k stars), Brightdata SDK JS (brightdata/skills, 264 stars), Anti Detect Browser (antibrow/anti-detect-browser-skills, 17 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.