Agent skill

Apify Core Workflow A

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Build a complete web scraping Actor with Crawlee and deploy to Apify.

MITAuto-check passedData & Analytics

Install Apify Core Workflow A

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-a -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace apify-core-workflow-a --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/apify-core-workflow-a .claude/skills/apify-core-workflow-a && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
apify-core-workflow-a
GitHub stars
2.8k
Token cost
~1.7k tokens
SKILL.md length
457 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Build a complete web scraping Actor with Crawlee and deploy to Apify.

  • Works in 6 steps: Define Input Schema → Build the Actor with Router Pattern → Configure Dockerfile → …
  • You need end-to-end web scraping on Apify: defining an input schema
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 4 more sections
  • Calls npm; reaches example-store.com; needs APIFY_TOKEN

What it does

Apify Core Workflow A is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build a complete web scraping Actor with Crawlee and deploy to Apify. Use when you need end-to-end web scraping on Apify: defining an input schema, building a router-based Crawlee crawler, extracting structured data, storing results in a dataset, testing locally, and deploying the Actor to the platform. Trigger with "apify scrape website", "build apify actor", "crawlee scraper", "apify main workflow".

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/examples.md` and `references/implementation.md`). Compatibility notes: Designed for Claude Code

It sits in Data & Analytics, covering Web scraping. It works with Apify and Docker. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • You need end-to-end web scraping on Apify: defining an input schema
  • Building a router-based Crawlee crawler
  • Extracting structured data
  • Storing results in a dataset

Example prompts

  • “apify scrape website”
  • “build apify actor”
  • “crawlee scraper”
  • “/apify-core-workflow-a”

Requirements

  • Node.js
  • A credential in APIFY_TOKEN
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(npm:*), Bash(npx:*), Bash(apify:*), Grep

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Define Input Schema
  2. Build the Actor with Router Pattern
  3. Configure Dockerfile
  4. Test Locally
  5. Deploy to Apify Platform
  6. Retrieve Results Programmatically

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(npm:*)
    • Bash(npx:*)
    • Bash(apify:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • example-store.com

    Also links to:

    • docs.apify.com
    • crawlee.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • APIFY_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Apify Core Workflow A loads about 1.7k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 457 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 457 words, ~1,732 tokens.

Download SKILL.mdSave it as .claude/skills/apify-core-workflow-a/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
apify-core-workflow-a
description
Build a complete web scraping Actor with Crawlee and deploy to Apify. Use when you need end-to-end web scraping on Apify: defining an input schema, building a router-based Crawlee crawler, extracting structured data, storing results in a dataset, testing locally, and deploying the Actor to the platform. Trigger with "apify scrape website", "build apify actor", "crawlee scraper", "apify main workflow".
allowed-tools
Read, Write, Edit, Bash(npm:*), Bash(npx:*), Bash(apify:*), Grep
compatibility
Designed for Claude Code
version
1.5.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, scraping, automation, apify

Apify Core Workflow A — Build & Deploy a Scraper

Overview

End-to-end workflow: define input schema, build a Crawlee-based Actor, extract structured data, store results in datasets, test locally, and deploy to Apify platform. This is the primary money-path workflow for Apify.

Prerequisites

  • npm install apify crawlee in your project
  • npm install -g apify-cli and apify login completed
  • For programmatic retrieval (Step 6), an API token in APIFY_TOKEN — read it from the environment (process.env.APIFY_TOKEN), never hard-code it
  • Familiarity with apify-sdk-patterns

Instructions

Step 1: Define Input Schema

Create .actor/INPUT_SCHEMA.json:

json
{
  "title": "E-Commerce Scraper",
  "type": "object",
  "schemaVersion": 1,
  "properties": {
    "startUrls": {
      "title": "Start URLs",
      "type": "array",
      "description": "Product listing page URLs to scrape",
      "editor": "requestListSources",
      "prefill": [{ "url": "https://example-store.com/products" }]
    },
    "maxItems": {
      "title": "Max items",
      "type": "integer",
      "description": "Maximum number of products to scrape",
      "default": 100,
      "minimum": 1,
      "maximum": 10000
    },
    "proxyConfig": {
      "title": "Proxy configuration",
      "type": "object",
      "description": "Select proxy to use",
      "editor": "proxy",
      "default": { "useApifyProxy": true }
    }
  },
  "required": ["startUrls"]
}
Step 2: Build the Actor with Router Pattern

Use a Crawlee router that splits handling by page type: the default handler enqueues product links + pagination from listing pages, and a PRODUCT-labeled handler extracts structured fields from detail pages. The entry point wires proxy config, concurrency, a failed-request handler, and a run summary into the key-value store. Skeleton:

typescript
// src/main.ts
import { Actor } from 'apify';
import { CheerioCrawler, createCheerioRouter, Dataset, log } from 'crawlee';

const router = createCheerioRouter();
router.addDefaultHandler(async ({ enqueueLinks }) => {
  await enqueueLinks({ selector: 'a.product-card', label: 'PRODUCT' });
  await enqueueLinks({ selector: 'a.next-page', label: 'LISTING' });
});
router.addHandler('PRODUCT', async ({ request, $ }) => {
  await Actor.pushData({ url: request.url, name: $('h1.product-title').text().trim() });
});

await Actor.main(async () => {
  const input = await Actor.getInput();
  const crawler = new CheerioCrawler({ requestHandler: router, maxRequestsPerCrawl: input?.maxItems ?? 100 });
  await crawler.run(input.startUrls.map(s => s.url));
});

The full typed Actor — Product/ProductInput interfaces, proxy configuration, failedRequestHandler, and the SUMMARY key-value write — is in implementation.md, Step 2.

Step 3: Configure Dockerfile

Use the apify/actor-node:20 base with a two-stage build (compile TypeScript in a builder stage, ship only dist/ + production deps). Full Dockerfile: implementation.md, Step 3.

Step 4: Test Locally
bash
# Create test input
mkdir -p storage/key_value_stores/default
echo '{"startUrls":[{"url":"https://example.com"}],"maxItems":5}' \
  > storage/key_value_stores/default/INPUT.json

# Run locally
apify run

# Check results
ls storage/datasets/default/
cat storage/key_value_stores/default/SUMMARY.json
Step 5: Deploy to Apify Platform
bash
# Push to Apify (creates Actor if it doesn't exist)
apify push

# Or push to a specific Actor
apify push username/my-actor

# Run on platform
apify actors call username/my-actor
Step 6: Retrieve Results Programmatically

From any client, use the apify-client SDK to call the deployed Actor, list its dataset items, and download results (JSON/CSV). The token comes from process.env.APIFY_TOKEN — never hard-code it. Full retrieval code: implementation.md, Step 6.

Output

  • Deployable Actor with typed input schema
  • Router-based crawler handling listing + detail pages
  • Structured product data in default dataset
  • Run summary in default key-value store
  • Failed requests tracked with error messages
Show full SKILL.md (187 more words)Show less

Error Handling

ErrorCauseSolution
Actor build failedDockerfile/deps issueCheck build logs on platform
Selector returns emptyPage structure changedUpdate CSS selectors
maxRequestsPerCrawl hitToo many pages enqueuedIncrease limit or filter URLs
Proxy errorsAnti-bot blockingSwitch to residential proxy
TIMED-OUT statusActor exceeded timeoutIncrease timeout or reduce scope

Examples

A quick example — seed a local input, run the Actor, and check results:

bash
mkdir -p storage/key_value_stores/default
echo '{"startUrls":[{"url":"https://example-store.com/products"}],"maxItems":5}' \
  > storage/key_value_stores/default/INPUT.json
apify run
cat storage/key_value_stores/default/SUMMARY.json

Three fuller worked scenarios live in examples.md:

  • Scrape a catalog locally, then deploy — the full seed → apify run → inspect → apify push loop, with the expected SUMMARY.json output.
  • Run the deployed Actor and export CSV — call the Actor via apify-client and download the dataset as CSV.
  • Route through residential proxy — pass a proxyConfig group at run time to get past anti-bot blocking.

Resources

Next Steps

Once your Actor is deployed and producing data, move on to dataset and key-value store management — pagination over large datasets, deduplication, exporting to external stores, and scheduling recurring runs — covered in apify-core-workflow-b.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/apify-core-workflow-a of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/examples.md
  • references/implementation.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Apify Core Workflow A next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Apify Core Workflow A compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Apify Core Workflow A this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.7kAutomated safety check: PassMIT
Apify Actorizationsickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
Apify Actor Developmentapify/agent-skills2.4k—~2.9kAutomated safety check: PassNone
Google Maps ScraperMahanaicoach/google-maps-scraper-kit1.3k—~2.8kAutomated safety check: PassMIT
Reddit Post Findergooseworks-ai/goose-skills1.2k1 repos~1.2kAutomated safety check: PassMIT
Apify CLIapify/apify-cli256—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • Apify Actorization

    sickn33/agentic-awesome-skills

    Actorization converts existing software into reusable serverless applications compatible with the Apify platform.

    47k GitHub starsUsed in 2 repos~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Apify Actor Development

    apify/agent-skills

    Official

    Creates, changes, debugs and deploys Apify Actors, including their input and output schemas, using the Apify CLI.

    2.4k GitHub stars~2.9k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Google Maps Scraper

    Mahanaicoach/google-maps-scraper-kit

    Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.

    1.3k GitHub stars~2.8k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Reddit Post Finder

    gooseworks-ai/goose-skills

    Scrape and search Reddit posts using Apify. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Data & AnalyticsAuto-check passed
  • Apify CLI

    apify/apify-cli

    Official

    Patterns for invoking the Apify CLI (apify) from agents. An agent skill from apify/apify-cli.

    256 GitHub stars~1.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Apify Collect

    extrasmall0/dear-hiring-manager

    Collect fresh job-posting URLs into ~/.dear-hiring-manager/urls.txt by running an Apify job scraper — the discovery source for /batch.

    111 GitHub stars~1.1k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Apify Core Workflow A

What does Apify Core Workflow A do?

Build a complete web scraping Actor with Crawlee and deploy to Apify. Apify Core Workflow A is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build a complete web scraping Actor with Crawlee and deploy to Apify.

When should I use Apify Core Workflow A?

Apify Core Workflow A fits situations like: you need end-to-end web scraping on Apify: defining an input schema; building a router-based Crawlee crawler; extracting structured data; storing results in a dataset.

How do I install Apify Core Workflow A in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-a -a claude-code`. Or copy the skill folder (skills/.curated/apify-core-workflow-a in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/apify-core-workflow-a in your project. Claude Code loads it when a task matches its description.

How do I install Apify Core Workflow A in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-a -a codex`. Or copy the skill folder (skills/.curated/apify-core-workflow-a in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/apify-core-workflow-a in your project. Codex loads it when a task matches its description.

Can I use Apify Core Workflow A in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-a -a cursor` (or -a -a, -a or -a for the others). To copy it by hand, put the folder in .cursor/skills/apify-core-workflow-a, .gemini/skills/apify-core-workflow-a, .github/skills/apify-core-workflow-a and .opencode/skills/apify-core-workflow-a in your project.

What does Apify Core Workflow A need to run?

Going by SKILL.md and its folder, Apify Core Workflow A needs the command-line tools its instructions call (npm) and credentials named APIFY_TOKEN. Our summary lists: Node.js; A credential in APIFY_TOKEN. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(npm:*), Bash(npx:*), Bash(apify:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Apify Core Workflow A access the network?

SKILL.md names 3 domains. In commands or code: example-store.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.apify.com and crawlee.dev. This is read from the text; nothing was executed.

Is Apify Core Workflow A safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Apify Core Workflow A use?

Apify Core Workflow A is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Apify Core Workflow A use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Apify Core Workflow A?

Skills that share tags, products or a category with Apify Core Workflow A: Apify Actorization (sickn33/agentic-awesome-skills, 47k stars), Apify Actor Development (apify/agent-skills, 2.4k stars), Google Maps Scraper (Mahanaicoach/google-maps-scraper-kit, 1.3k stars) and Reddit Post Finder (gooseworks-ai/goose-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Apify Core Workflow A?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.