Agent skill

Crawlee

by LeoYeAI in LeoYeAI/openclaw-master-skills

Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).

MITAuto-check passedData & Analytics

Install Crawlee

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills crawlee --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/crawlee .claude/skills/crawlee && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawlee
GitHub stars
2.2k
Token cost
~4.7k tokens
SKILL.md length
478 words
Files
4 (incl. references)
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).

  • Works in 12 steps: Choose Your Crawler → Installation → Core Concepts → …
  • The user wants to: scrape a website
  • SKILL.md covers 1. Choose Your Crawler, 2. Installation, 3. Core Concepts and 4. Quick Start Examples, plus 8 more sections
  • Calls npm, pip and npx; reaches proxy1.com and proxy2.com

What it does

Crawlee is an agent skill from LeoYeAI/openclaw-master-skills. Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python). Use this skill whenever the user wants to: scrape a website, build a web crawler, extract data from web pages, automate browser navigation, handle anti-bot blocking, manage proxies or sessions for scraping, use Playwright/Puppeteer/Cheerio/BeautifulSoup for web data extraction, crawl sitemaps, download files from URLs, or deploy a scraper to the cloud. Trigger even for loosely related phrases like "get data from…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `_meta.json`, `references/js-api.md` and `references/python-api.md`).

It sits in Data & Analytics, covering Web scraping. It works with Python, JavaScript, TypeScript and Playwright. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • The user wants to: scrape a website
  • Build a web crawler
  • Extract data from web pages
  • Automate browser navigation

Example prompts

  • “get data from a website”
  • “automate browser”
  • “scrape prices”
  • “/crawlee”

Requirements

  • Python 3
  • Node.js

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Choose Your Crawler
  2. Installation
  3. Core Concepts
  4. Quick Start Examples
  5. Routing — Handling Multiple Page Types
  6. Enqueuing Links
  7. Storage
  8. Proxy Management
  9. Session Management
  10. Avoiding Blocks
  11. Concurrency & Scaling
  12. Configuration & Environment Variables

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • pip
    • npx
    • playwright

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • proxy1.com
    • proxy2.com

    Also links to:

    • crawlee.dev
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawlee loads about 4.7k tokens when it runs, and up to ~8.2k if it reads all its reference files. Until then it costs about 198 tokens; SKILL.md has 478 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~198
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 478 words, ~4,725 tokens.

Download SKILL.mdSave it as .claude/skills/crawlee/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
crawlee
description
Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python). Use this skill whenever the user wants to: scrape a website, build a web crawler, extract data from web pages, automate browser navigation, handle anti-bot blocking, manage proxies or sessions for scraping, use Playwright/Puppeteer/Cheerio/BeautifulSoup for web data extraction, crawl sitemaps, download files from URLs, or deploy a scraper to the cloud. Trigger even for loosely related phrases like "get data from a website", "automate browser", "scrape prices", "extract links", "crawl URLs", or "bypass bot detection". Covers CheerioCrawler, PlaywrightCrawler, PuppeteerCrawler, HttpCrawler, JSDOMCrawler (JS), and BeautifulSoupCrawler, ParselCrawler, PlaywrightCrawler (Python).

Crawlee Skill

Crawlee is a production-grade web scraping and browser automation library for JavaScript/TypeScript (Node.js 16+) and Python (3.10+). It handles anti-blocking, proxies, session management, storage, and concurrency out of the box.

Docs: https://crawlee.dev/js/docs | https://crawlee.dev/python/docs
GitHub: https://github.com/apify/crawlee


1. Choose Your Crawler

JavaScript / TypeScript
CrawlerWhen to UseJS Required
CheerioCrawlerFast HTML parsing, no JS rendering needed❌
HttpCrawlerRaw HTTP responses, custom parsing❌
JSDOMCrawlerDOM manipulation without full browser❌
PlaywrightCrawlerModern headless browser (Chromium/Firefox/WebKit)✅
PuppeteerCrawlerChromium/Chrome headless automation✅
AdaptivePlaywrightCrawlerAuto-detects if JS rendering is neededAuto
BasicCrawlerCustom HTTP logic from scratch❌

Rule of thumb: Start with CheerioCrawler. Upgrade to PlaywrightCrawler only when JS rendering is required.

Python
CrawlerWhen to Use
BeautifulSoupCrawlerHTML parsing with BeautifulSoup (fast, no JS)
ParselCrawlerCSS/XPath selectors, Scrapy-style (fast, no JS)
PlaywrightCrawlerFull browser automation (Chromium/Firefox/WebKit)
AdaptivePlaywrightCrawlerAuto HTTP vs browser decision

2. Installation

JavaScript
bash
# Recommended: use the CLI
npx crawlee create my-crawler
cd my-crawler && npm install

# Or manually:
npm install crawlee

# For Playwright:
npm install crawlee playwright
npx playwright install

# For Puppeteer:
npm install crawlee puppeteer

Add to package.json:

json
{ "type": "module" }
Python
bash
pip install crawlee

# With BeautifulSoup:
pip install 'crawlee[beautifulsoup]'

# With Playwright:
pip install 'crawlee[playwright]'
playwright install

3. Core Concepts

The Two Questions Every Crawler Answers
  1. Where to go? → Request objects in a RequestQueue
  2. What to do there? → requestHandler function (JS) / decorated handler (Python)
Key Classes (JS)
  • Request — A single URL + metadata to crawl
  • RequestQueue — Dynamic, deduplicated queue of URLs
  • Dataset — Append-only structured result storage (like a table)
  • KeyValueStore — Blob storage for screenshots, PDFs, state
  • ProxyConfiguration — Manages proxy rotation
  • SessionPool — Manages browser sessions + cookies

4. Quick Start Examples

javascript
import { CheerioCrawler, Dataset } from 'crawlee';

const crawler = new CheerioCrawler({
  async requestHandler({ $, request, enqueueLinks, log }) {
    const title = $('title').text();
    log.info(`Title of ${request.loadedUrl}: ${title}`);

    await Dataset.pushData({ url: request.loadedUrl, title });

    // Enqueue all links found on this page
    await enqueueLinks();
  },
  maxRequestsPerCrawl: 100, // Safety limit
});

await crawler.run(['https://example.com']);
JavaScript — PlaywrightCrawler
javascript
import { PlaywrightCrawler, Dataset } from 'crawlee';

const crawler = new PlaywrightCrawler({
  // headless: false, // Uncomment to see the browser
  async requestHandler({ page, request, enqueueLinks, log }) {
    const title = await page.title();
    log.info(`${request.loadedUrl}: ${title}`);
    await Dataset.pushData({ url: request.loadedUrl, title });
    await enqueueLinks();
  },
});

await crawler.run(['https://example.com']);
Python — BeautifulSoupCrawler
python
import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext

async def main() -> None:
    crawler = BeautifulSoupCrawler(max_requests_per_crawl=50)

    @crawler.router.default_handler
    async def handler(context: BeautifulSoupCrawlingContext) -> None:
        title = context.soup.title.string if context.soup.title else None
        context.log.info(f'Processing {context.request.url}: {title}')
        await context.push_data({'url': context.request.url, 'title': title})
        await context.enqueue_links()

    await crawler.run(['https://example.com'])

if __name__ == '__main__':
    asyncio.run(main())
Python — PlaywrightCrawler
python
import asyncio
from crawlee.crawlers import PlaywrightCrawler, PlaywrightCrawlingContext

async def main() -> None:
    crawler = PlaywrightCrawler(headless=True, browser_type='chromium')

    @crawler.router.default_handler
    async def handler(context: PlaywrightCrawlingContext) -> None:
        title = await context.page.title()
        await context.push_data({'url': context.request.url, 'title': title})
        await context.enqueue_links()

    await crawler.run(['https://example.com'])

if __name__ == '__main__':
    asyncio.run(main())

5. Routing — Handling Multiple Page Types

Use labels + router to handle different kinds of pages (list pages, detail pages, etc.).

JavaScript
javascript
import { PlaywrightCrawler, Dataset } from 'crawlee';
import { router } from './routes.js';

const crawler = new PlaywrightCrawler({ requestHandler: router });

await crawler.run([{ url: 'https://shop.example.com', label: 'START' }]);
javascript
// routes.js
import { createPlaywrightRouter } from 'crawlee';

export const router = createPlaywrightRouter();

router.addHandler('START', async ({ page, enqueueLinks }) => {
  await enqueueLinks({ selector: 'a.category', label: 'CATEGORY' });
});

router.addHandler('CATEGORY', async ({ page, enqueueLinks }) => {
  await enqueueLinks({ selector: 'a.product', label: 'DETAIL' });
  // Enqueue next page
  const next = await page.$('a.next-page');
  if (next) await enqueueLinks({ selector: 'a.next-page', label: 'CATEGORY' });
});

router.addDefaultHandler(async ({ page, request, pushData }) => {
  // DETAIL pages
  const title = await page.title();
  const price = await page.$eval('.price', el => el.textContent);
  await pushData({ url: request.url, title, price });
});
Python
python
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext

crawler = BeautifulSoupCrawler()

@crawler.router.handler('CATEGORY')
async def category_handler(context: BeautifulSoupCrawlingContext) -> None:
    await context.enqueue_links(selector='a.product', label='DETAIL')

@crawler.router.default_handler
async def detail_handler(context: BeautifulSoupCrawlingContext) -> None:
    title = context.soup.title.string
    await context.push_data({'url': context.request.url, 'title': title})

javascript
// Enqueue all links on page
await enqueueLinks();

// Filter by glob pattern
await enqueueLinks({ globs: ['https://example.com/products/**'] });

// Filter by regex
await enqueueLinks({ regexps: [/\/product\/\d+/] });

// Enqueue only specific selector
await enqueueLinks({ selector: 'a.pagination', label: 'LIST' });

// Enqueue with custom label and transformations
await enqueueLinks({
  selector: 'a.item',
  label: 'DETAIL',
  transformRequestFunction: (req) => {
    req.userData.scrapedAt = new Date().toISOString();
    return req;
  },
});
Python
python
await context.enqueue_links()
await context.enqueue_links(selector='a.product', label='DETAIL')
await context.enqueue_links(include=[re.compile(r'/products/\d+')])

7. Storage

Dataset (structured results)
javascript
// JS — Write
await Dataset.pushData({ url, title, price });
await Dataset.pushData([item1, item2, item3]); // batch write

// JS — Read / Export
const dataset = await Dataset.open();
await dataset.exportToCSV('results'); // saves to KV store
await dataset.exportToJSON('results');

for await (const item of dataset) { console.log(item); }
python
# Python — Write
await context.push_data({'url': url, 'title': title})

# Python — Read / Export
from crawlee.storages import Dataset
dataset = await Dataset.open()
await dataset.export_to(key='results', content_type='csv')

Data is saved to ./storage/datasets/default/*.json by default.

KeyValueStore (blobs, screenshots, state)
javascript
// JS
await KeyValueStore.setValue('OUTPUT', { results: [...] });
const value = await KeyValueStore.getValue('OUTPUT');

// Save a screenshot
const store = await KeyValueStore.open();
await store.setValue('screenshot', await page.screenshot(), { contentType: 'image/png' });
python
# Python
from crawlee.storages import KeyValueStore
kvs = await KeyValueStore.open()
await kvs.set_value('result', {'data': 'value'})
value = await kvs.get_value('result')
Storage location
./storage/
  datasets/default/     # Dataset rows as JSON files
  key_value_stores/default/  # KV store entries
  request_queues/default/    # Request queue state

Override with env var: CRAWLEE_STORAGE_DIR=/path/to/storage


8. Proxy Management

javascript
// JS — Basic proxy rotation
import { ProxyConfiguration } from 'crawlee';

const proxyConfiguration = new ProxyConfiguration({
  proxyUrls: [
    'http://user:pass@proxy1.example.com:8000',
    'http://user:pass@proxy2.example.com:8000',
  ],
});

const crawler = new CheerioCrawler({
  proxyConfiguration,
  useSessionPool: true,
  persistCookiesPerSession: true,
  async requestHandler({ proxyInfo, request }) {
    console.log('Using proxy:', proxyInfo?.url);
  },
});
javascript
// JS — Tiered proxies (smart cost/reliability balancing)
const proxyConfiguration = new ProxyConfiguration({
  tieredProxyUrls: [
    [null],                              // Tier 0: no proxy (cheapest)
    ['http://cheap-datacenter-proxy'],   // Tier 1: datacenter
    ['http://expensive-residential'],    // Tier 2: residential (most reliable)
  ],
});
// Crawlee auto-escalates tiers when blocking is detected, then drops back when clear
python
# Python
from crawlee.proxy_configuration import ProxyConfiguration

proxy_configuration = ProxyConfiguration(
    proxy_urls=['http://proxy1.com/', 'http://proxy2.com/'],
)
crawler = BeautifulSoupCrawler(
    proxy_configuration=proxy_configuration,
    use_session_pool=True,
)

Show full SKILL.md (196 more words)Show less

9. Session Management

Sessions tie together cookies, proxy IPs, and headers to simulate a consistent user identity.

javascript
// JS
const crawler = new CheerioCrawler({
  useSessionPool: true,         // Enable (default: true)
  persistCookiesPerSession: true,
  sessionPoolOptions: { maxPoolSize: 100 },

  async requestHandler({ session, $ }) {
    const title = $('title').text();
    if (title === 'Access Denied') {
      session?.retire();  // Mark this IP+cookie combo as blocked
    } else if (title === 'Slow') {
      session?.markBad(); // Penalize but don't retire
    }
    // session.markGood() is called automatically on success
  },
});
python
# Python
from crawlee.sessions import SessionPool

crawler = BeautifulSoupCrawler(
    use_session_pool=True,
    session_pool=SessionPool(max_pool_size=100),
)

@crawler.router.default_handler
async def handler(context: BeautifulSoupCrawlingContext) -> None:
    title = context.soup.title.string if context.soup.title else ''
    if title == 'Access Denied':
        context.session.retire()

10. Avoiding Blocks

javascript
// JS — Playwright with fingerprint rotation (built-in, zero config needed)
const crawler = new PlaywrightCrawler({
  // Fingerprints automatically randomized by default in Playwright/Puppeteer crawlers
  // headless: false,  // Use headful for harder targets
  async requestHandler({ page }) {
    // Add realistic delays
    await page.waitForTimeout(1000 + Math.random() * 2000);
  },
});

// Use got-scraping for HTTP (built into CheerioCrawler/HttpCrawler)
// It automatically sets realistic headers and TLS fingerprints

Anti-blocking checklist:

  • ✅ Use CheerioCrawler — it uses got-scraping which mimics real browser HTTP
  • ✅ Enable useSessionPool: true with a proxyConfiguration
  • ✅ Use tiered proxies for automatic failover
  • ✅ Set maxRequestsPerMinute to avoid rate limits
  • ✅ For browser crawlers — fingerprints are rotated automatically
  • ✅ Use persistCookiesPerSession: true
  • ✅ Retire sessions on blocks: session.retire()

11. Concurrency & Scaling

javascript
// JS
const crawler = new CheerioCrawler({
  maxConcurrency: 50,         // Max parallel requests (default: 200)
  minConcurrency: 1,          // Don't set too high!
  maxRequestsPerMinute: 120,  // Rate limit
  maxRequestsPerCrawl: 1000,  // Total request cap (safety)
  requestHandlerTimeoutSecs: 30,
});
python
# Python
from crawlee import ConcurrencySettings

crawler = BeautifulSoupCrawler(
    concurrency_settings=ConcurrencySettings(
        max_concurrency=50,
        max_tasks_per_minute=120,
    ),
    max_requests_per_crawl=1000,
)

Scaling notes:

  • Crawlee auto-scales concurrency based on CPU/memory
  • Don't set minConcurrency high — it can crash under load
  • maxRequestsPerMinute is smoother than raw concurrency throttling

12. Configuration & Environment Variables

Env VariableDefaultPurpose
CRAWLEE_STORAGE_DIR./storageStorage root directory
CRAWLEE_DEFAULT_DATASET_IDdefaultOverride default dataset ID
CRAWLEE_DEFAULT_KEY_VALUE_STORE_IDdefaultOverride default KVS ID
CRAWLEE_DEFAULT_REQUEST_QUEUE_IDdefaultOverride default queue ID
CRAWLEE_PURGE_ON_STARTtrueClear storage before each run
javascript
// JS — Programmatic configuration
import { Configuration } from 'crawlee';

const config = new Configuration({
  storageDir: '/data/crawlee',
  persistStateIntervalMillis: 30_000,
});

const crawler = new CheerioCrawler({ /* ... */ }, config);

13. Docker Deployment

dockerfile
FROM apify/actor-node-playwright-chrome:20

COPY package*.json ./
RUN npm ci --only=prod

COPY . ./

CMD ["node", "src/main.js"]

For Cheerio (smaller image):

dockerfile
FROM apify/actor-node:20

14. Common Patterns

Pagination
javascript
// JS — Enqueue next page
router.addHandler('LIST', async ({ page, enqueueLinks }) => {
  await enqueueLinks({ selector: '.product', label: 'DETAIL' });
  const hasNext = await page.$('a.next');
  if (hasNext) await enqueueLinks({ selector: 'a.next', label: 'LIST' });
});
Downloading Files
javascript
// JS — Save to KeyValueStore
const { body } = await sendRequest({ responseType: 'buffer' });
await KeyValueStore.setValue('file.pdf', body, { contentType: 'application/pdf' });
Taking Screenshots
javascript
// JS — Playwright
async requestHandler({ page, request }) {
  const screenshot = await page.screenshot({ fullPage: true });
  await KeyValueStore.setValue(
    `screenshot-${Date.now()}`,
    screenshot,
    { contentType: 'image/png' }
  );
}
Shared State Across Handlers
javascript
// JS — useState()
async requestHandler({ useState }) {
  const state = await useState({ count: 0 });
  state.count++;
  console.log('Total processed:', state.count);
}
Error Handling & Retries
javascript
// JS
const crawler = new CheerioCrawler({
  maxRequestRetries: 3, // Retry failed requests up to 3 times
  failedRequestHandler: async ({ request, error }) => {
    console.error(`Failed: ${request.url}`, error.message);
    await Dataset.pushData({ url: request.url, error: error.message });
  },
});
python
# Python
crawler = BeautifulSoupCrawler(max_request_retries=3)

@crawler.failed_request_handler
async def on_failed(context: BasicCrawlingContext, error: Exception) -> None:
    context.log.error(f'Failed {context.request.url}: {error}')
Sitemap Crawling
javascript
import { CheerioCrawler } from 'crawlee';
import { Sitemap } from '@crawlee/utils';

const { urls } = await Sitemap.load('https://example.com/sitemap.xml');
const crawler = new CheerioCrawler({ /* ... */ });
await crawler.run(urls);
Run as Web Server
javascript
import { CheerioCrawler } from 'crawlee';
import { createServer } from 'http';

const server = createServer(async (req, res) => {
  const url = new URL(req.url, 'http://localhost').searchParams.get('url');
  const crawler = new CheerioCrawler({
    maxRequestsPerCrawl: 1,
    async requestHandler({ $ }) {
      res.end(JSON.stringify({ title: $('title').text() }));
    },
  });
  await crawler.run([url]);
});
server.listen(3000);

15. TypeScript Support

typescript
import { CheerioCrawler, CheerioCrawlingContext, Dataset } from 'crawlee';

interface Product {
  url: string;
  title: string;
  price: number;
}

const crawler = new CheerioCrawler({
  async requestHandler({ $, request }: CheerioCrawlingContext) {
    const title = $('h1').text();
    const price = parseFloat($('.price').text().replace('$', ''));
    await Dataset.pushData<Product>({ url: request.url, title, price });
  },
});

16. Cloud Deployment (Apify Platform)

javascript
import { Actor } from 'apify';
import { CheerioCrawler } from 'crawlee';

await Actor.init();

const input = await Actor.getInput();
const { startUrls } = input;

const crawler = new CheerioCrawler({
  async requestHandler({ $, request }) {
    await Actor.pushData({ url: request.url, title: $('title').text() });
  },
});

await crawler.run(startUrls);
await Actor.exit();

Deploy with: apify push


17. Debugging Tips

javascript
// Enable verbose logging
import { Log } from 'crawlee';
Log.setLevel(Log.LEVELS.DEBUG);

// Run headful (browser crawlers only)
const crawler = new PlaywrightCrawler({
  headless: false,
  // ...
});

// Limit requests while developing
const crawler = new CheerioCrawler({
  maxRequestsPerCrawl: 10,
  // ...
});

18. Reference Files

For advanced topics, see:

  • references/js-api.md — Full JS API quick reference
  • references/python-api.md — Full Python API quick reference

Both language docs: https://crawlee.dev

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/crawlee of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json
  • references/js-api.md
  • references/python-api.md

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Crawlee next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawlee compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawlee this skillLeoYeAI/openclaw-master-skills2.2k—~4.7kAutomated safety check: PassMIT
Apify Actor Developmentapify/agent-skills2.4k—~2.9kAutomated safety check: PassNone
Brightdata SDK JSbrightdata/skills264—~3kAutomated safety check: PassMIT
Anti Detect Browserantibrow/anti-detect-browser-skills17—~9.8kAutomated safety check: WarnMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Agent Browseroxylabs/agent-skills875—~3kAutomated safety check: PassMIT

Similar skills

  • Apify Actor Development

    apify/agent-skills

    Official

    Creates, changes, debugs and deploys Apify Actors, including their input and output schemas, using the Apify CLI.

    2.4k GitHub stars~2.9k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    17 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Agent Browser

    oxylabs/agent-skills

    Connects to Oxylabs remote agent browsers over the Chrome DevTools Protocol (CDP) with Playwright or Puppeteer.

    875 GitHub stars~3k tokensUpdated 8 days ago
    Productivity & AutomationAuto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Crawlee

What does Crawlee do?

Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python). Crawlee is an agent skill from LeoYeAI/openclaw-master-skills. Expert guide for building web scrapers and crawlers using Crawlee (JavaScript/TypeScript and Python).

When should I use Crawlee?

Crawlee fits situations like: the user wants to: scrape a website; build a web crawler; extract data from web pages; automate browser navigation.

How do I install Crawlee in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a claude-code`. Or copy the skill folder (skills/crawlee in LeoYeAI/openclaw-master-skills) into .claude/skills/crawlee in your project. Claude Code loads it when a task matches its description.

How do I install Crawlee in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a codex`. Or copy the skill folder (skills/crawlee in LeoYeAI/openclaw-master-skills) into .agents/skills/crawlee in your project. Codex loads it when a task matches its description.

Can I use Crawlee in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill crawlee -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawlee, .gemini/skills/crawlee, .github/skills/crawlee and .opencode/skills/crawlee in your project.

What does Crawlee need to run?

Going by SKILL.md and its folder, Crawlee needs the command-line tools its instructions call (npm, pip, npx and playwright). Our summary lists: Python 3; Node.js.

Does Crawlee access the network?

SKILL.md names 4 domains. In commands or code: proxy1.com and proxy2.com; the agent is likely to contact these when it follows the instructions. As links in the text: crawlee.dev and github.com. This is read from the text; nothing was executed.

Is Crawlee safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Crawlee use?

Crawlee is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawlee use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Crawlee?

Skills that share tags, products or a category with Crawlee: Apify Actor Development (apify/agent-skills, 2.4k stars), Brightdata SDK JS (brightdata/skills, 264 stars), Anti Detect Browser (antibrow/anti-detect-browser-skills, 17 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawlee?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.