Agent skill

Extract

by actionbook in actionbook/actionbook

Extract structured data from websites and produce an executable Playwright script plus extracted data.

Apache-2.0Auto-check passedSales & Support

Install Extract

skills CLI
$ npx skills add actionbook/actionbook --skill extract -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install actionbook/actionbook extract --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/actionbook/actionbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/extract .claude/skills/extract && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extract
GitHub stars
1.6k
Used in
2 other repos
Token cost
~3.4k tokens
SKILL.md length
1,005 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract structured data from websites and produce an executable Playwright script plus extracted data.

  • Works in 6 steps: Understand the target → Obtain selectors and choose execution path → Probe page mechanisms and fallback only… → …
  • The user wants to scrape
  • SKILL.md covers When to Use This Skill, Decision Strategy, Mechanism-Aware Script Strategy and Execution Chain, plus 3 more sections
  • Calls node

What it does

Extract is an agent skill from actionbook/actionbook. Extract structured data from websites and produce an executable Playwright script plus extracted data. Use when the user wants to scrape, extract, pull, collect, or harvest data from any website — product listings, tables, search results, feeds, profiles, or any repeating content.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Sales & Support, covering Web scraping, Schema markup and E-commerce operations. It works with Playwright. The repository describes itself as: Let your AI agent get the sources behind logins and paywalls. The licence is Apache-2.0.

When your agent uses it

  • The user wants to scrape
  • Harvest data from any website — product listings
  • Any repeating content

Example prompts

  • “/extract”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Understand the target
  2. Obtain selectors and choose execution path
  3. Probe page mechanisms and fallback only when needed
  4. Generate Playwright script
  5. Execute and validate
  6. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit 0e31254. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extract loads about 3.4k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 1,005 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from actionbook/actionbook at commit 0e31254, republished under its Apache-2.0 licence (© actionbook). 1,005 words, ~3,375 tokens.

Download SKILL.mdSave it as .claude/skills/extract/SKILL.md (or your agent's skills folder).
name
extract
description
Extract structured data from websites and produce an executable Playwright script plus extracted data. Use when the user wants to scrape, extract, pull, collect, or harvest data from any website — product listings, tables, search results, feeds, profiles, or any repeating content.

When to Use This Skill

Activate when the user wants to obtain data from a website:

  • "Extract all product prices from this page"
  • "Scrape the table of results from ..."
  • "Pull the list of authors and titles from arXiv search results"
  • "Collect all job listings from this page"
  • "Get the data from this dashboard table"
  • "Harvest review scores from ..."
  • "Download all the links/images/cards from ..."

The deliverable is always two artifacts:

  1. Executable Playwright script — a standalone .cjs file that reproduces the extraction without Actionbook at runtime.
  2. Extracted data — JSON (default), CSV, or user-specified format written to disk.

Decision Strategy

Use Actionbook as a conditional accelerator, not a mandatory step. The goal is reliable selectors in the shortest path.

User request
  │
  ├─► actionbook search "<site> <intent>"
  │     ├─ Results with Health Score ≥ 70%  ──► actionbook get "<ID>" ──► use selectors
  │     └─ No results / low score  ──► Fallback
  │
  └─► Fallback: actionbook browser open <url>
        ├─ actionbook browser snapshot   (accessibility tree → find selectors)
        ├─ actionbook browser screenshot (visual confirmation)
        └─ manual selector discovery via DOM inspection

Priority order for selector sources:

PrioritySourceWhen
1actionbook getSite is indexed, health score ≥ 70%
2actionbook browser snapshotNot indexed or selectors outdated
3DOM inspection via screenshot + snapshotComplex SPA / dynamic content

Non-negotiable rule: if search + get already provides usable selectors for required fields, start from get selectors and do not jump to full fallback (snapshot/screenshot) by default. Exception: lightweight mechanism probes (for hydration/virtualization/pagination) are allowed when runtime behavior may affect script correctness. Escalate to snapshot/screenshot only when probes/sample validation indicate selector gaps or instability.

Mechanism-Aware Script Strategy

Websites use patterns that break naive scraping. The generated Playwright script must account for these:

Streaming / SSR / RSC hydration

Pages may render a shell first, then stream or hydrate content.

javascript
// Wait for hydration to complete — not just DOMContentLoaded
await page.waitForSelector('[data-item]', { state: 'attached' });
await page.waitForFunction(() => {
  const items = document.querySelectorAll('[data-item]');
  return items.length > 0 && !document.querySelector('[data-pending]');
});

Detection cues: React root with data-reactroot, Next.js __NEXT_DATA__, empty containers that fill after JS runs. If actionbook browser text "<selector>" returns empty but the screenshot shows content, hydration hasn't completed.

Virtualized lists / virtual DOM

Only visible rows exist in the DOM. Scrolling renders new rows and destroys old ones.

javascript
// Scroll-and-collect loop for virtualized lists (scroll container aware)
const allItems = [];
const maxScrolls = 50;
let scrolls = 0;

const container = await page.$('<scroll-container-selector>');
if (!container) throw new Error('Scroll container not found');

let previousTop = await container.evaluate(el => el.scrollTop);
while (scrolls < maxScrolls) {
  const items = await page.$$eval('[data-row]', rows =>
    rows.map(r => ({ text: r.textContent.trim() }))
  );
  for (const item of items) {
    if (!allItems.find(i => i.text === item.text)) allItems.push(item);
  }

  await container.evaluate(el => el.scrollBy(0, 600));
  await page.waitForTimeout(300);

  const currentTop = await container.evaluate(el => el.scrollTop);
  if (currentTop === previousTop) break;

  previousTop = currentTop;
  scrolls += 1;
}

Detection cues: Container has fixed height with overflow: auto/scroll, row count in DOM is much smaller than stated total, rows have transform: translateY(...) or position: absolute; top: ...px.

Infinite scroll / lazy loading

New content appends when the user scrolls near the bottom.

javascript
// Scroll to bottom until no new content loads (with no-growth tolerance)
let itemCount = 0;
let noGrowthStreak = 0;
const maxScrolls = 80;
let scrolls = 0;

while (scrolls < maxScrolls && noGrowthStreak < 3) {
  await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
  await page.waitForTimeout(1200);

  const newCount = await page.$$eval('.item', els => els.length);
  if (newCount > itemCount) {
    itemCount = newCount;
    noGrowthStreak = 0;
  } else {
    noGrowthStreak += 1;
  }

  scrolls += 1;
}

Detection cues: Intersection Observer in page JS, "Load more" button, sentinel element at bottom, network requests firing on scroll.

Pagination

Multi-page results behind "Next" buttons or numbered pages.

javascript
// Click-through pagination (navigation-aware, SPA-safe)
const allData = [];
const maxPages = 50;
let pageIndex = 0;
while (pageIndex < maxPages) {
  const pageData = await page.$$eval('.result-item', items =>
    items.map(el => ({ title: el.querySelector('h3')?.textContent?.trim() }))
  );
  allData.push(...pageData);

  const nextBtn = await page.$('a.next-page:not([disabled])');
  if (!nextBtn) break;

  const previousUrl = page.url();
  const previousFirstItem = await page
    .$eval('.result-item', el => el.textContent?.trim() || '')
    .catch(() => '');

  await nextBtn.click();

  // Post-click detection only: advance must be caused by this click
  const advanced = await Promise.any([
    page
      .waitForURL(url => url.toString() !== previousUrl, { timeout: 5000 })
      .then(() => true),
    page
      .waitForFunction(
        prev => {
          const first = document.querySelector('.result-item');
          return !!first && (first.textContent || '').trim() !== prev;
        },
        previousFirstItem,
        { timeout: 5000 }
      )
      .then(() => true),
  ]).catch(() => false);

  if (!advanced) break;

  await page.waitForLoadState('networkidle').catch(() => {});
  pageIndex += 1;
}

Execution Chain

Step 1: Understand the target

Identify from the user request:

  • URL — the page to extract from
  • Data shape — what fields / columns are needed
  • Scope — single page, paginated, infinite scroll, or multi-page crawl
  • Output format — JSON (default), CSV, or other
Step 2: Obtain selectors and choose execution path
bash
# Try Actionbook index first
actionbook search "<site> <data-description>" --domain <domain>

# If good results (health ≥ 70%), get full selectors
actionbook get "<ID>"

Use this routing strictly:

  • Path A (default when get is good): requested fields are covered by get selectors and quality is acceptable.

    • Start from get selectors and move to script draft quickly.
    • You may run lightweight mechanism probes (browser text, quick scroll checks) before finalizing script strategy.
    • Do not run full fallback (snapshot / screenshot) before first draft unless probe/sample validation shows mismatch.
    • Field mapping must default to get selectors and mark source as actionbook_get.
  • Path B (partial / unstable): get exists but required fields are missing, selector resolves 0 elements, or validation fails.

    • Run targeted fallback only for failed fields/steps.
  • Path C (no usable coverage): search/get has no usable result.

    • Run full fallback discovery.
Step 3: Probe page mechanisms and fallback only when needed

Path A mechanism detection timing:

  • Run minimal probes either before final script draft or during sample validation.
  • Before any probe command, ensure the correct page context is open:
    • actionbook browser open "<url>" (if current tab context is unknown/stale)
  • If probes/sample run indicate mismatch (missing rows, unstable selectors, wrong pagination behavior), escalate to Path B targeted fallback.

Fallback discovery by path:

Path B targeted fallback (only failed fields/steps):

bash
actionbook browser open "<url>"     # if not already open
actionbook browser snapshot          # focus on failed field/container mapping
# actionbook browser screenshot      # optional visual confirmation for failed area

Path C full fallback (no usable coverage):

bash
actionbook browser open "<url>"
actionbook browser snapshot
actionbook browser screenshot

Mechanism probes (run when script strategy needs confirmation):

bash
# Hydration / streaming check
actionbook browser text "<container-selector>"

# Infinite scroll quick signal (explicit before/after decision)
actionbook browser eval "document.querySelectorAll('<item-selector>').length"   # before
actionbook browser click "<scroll-container-selector-or-body>"                    # focus scroll context
actionbook browser eval "const c=document.querySelector('<scroll-container-selector>') || document.scrollingElement; c.scrollBy(0, c.clientHeight || window.innerHeight);"
actionbook browser eval "document.querySelectorAll('<item-selector>').length"   # after
# If count increases, treat page as lazy-load/infinite-scroll.

Fallback trigger conditions:

  • actionbook get cannot map all required fields.
  • actionbook get selectors return empty/unstable values in sample run.
  • Runtime behavior conflicts with expected mechanism (e.g., virtualized container, delayed hydration).
Show full SKILL.md (362 more words)Show less
Step 4: Generate Playwright script

Write a standalone Playwright script (extract_<domain>_<slug>.cjs) that:

  1. Navigates to the target URL.
  2. Waits for the correct readiness signal (not just load — see mechanisms above).
  3. Handles the detected mechanism (virtual scroll, pagination, etc.).
  4. Extracts data into structured objects.
  5. Writes output to disk (JSON.stringify / CSV).
  6. Closes the browser.
  7. Enforces guardrails (maxPages, maxScrolls, timeout budget) to avoid infinite loops.

Script template:

javascript
// extract_<domain>_<slug>.cjs
// Generated by Actionbook extract skill
// Usage: node extract_<domain>_<slug>.cjs

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();

  await page.goto('<URL>', { waitUntil: 'domcontentloaded' });

  // -- wait for readiness --
  await page.waitForSelector('<container>', { state: 'visible' });

  // -- extract --
  const data = await page.$$eval('<item-selector>', items =>
    items.map(el => ({
      // fields mapped from user request
    }))
  );

  // -- output --
  const fs = require('fs');
  fs.writeFileSync('output.json', JSON.stringify(data, null, 2));
  console.log(`Extracted ${data.length} items → output.json`);

  await browser.close();
})();
Step 5: Execute and validate

Run the script to confirm it works:

bash
node extract_<domain>_<slug>.cjs

Validation rules:

CheckPass condition
Script exits 0No runtime errors
Output file existsNon-empty file written
Record count > 0At least one item extracted
No null/empty fieldsEvery declared field has a value in ≥ 90% of records
Data matches pageSpot-check first and last record against actionbook browser text

If validation fails, inspect the output, adjust selectors or wait strategy, and re-run.

Step 6: Deliver

Present to the user:

  1. Script path — the .cjs file they can re-run anytime.
  2. Data path — the output JSON/CSV file.
  3. Record count — how many items were extracted.
  4. Notes — any mechanism-specific caveats (e.g., "this site uses infinite scroll; the script scrolls up to 50 pages by default").

Output Contract

Every extract invocation produces:

ArtifactPathFormat
Playwright script./extract_<domain>_<slug>.cjsStandalone Node.js script using playwright
Extracted data./output.json (default) or user-specified pathJSON array of objects (default), CSV, or user-specified

The script must be re-runnable — a user should be able to execute it later without Actionbook installed, as long as Node.js + Playwright are available in the runtime environment.

Selector Priority

When multiple selector types are available from actionbook get:

PriorityTypeReason
1data-testidStable, test-oriented, rarely changes
2aria-labelAccessibility-driven, semantically meaningful
3CSS selectorStructural, may break on redesign
4XPathLast resort, most brittle

Error Handling

ErrorAction
actionbook search returns no resultsFall back to snapshot + screenshot
Selector returns 0 elementsRe-snapshot, compare with screenshot, update selector
Script times outAdd longer waitForTimeout, check for anti-bot measures
Partial data (some fields empty)Check if content is lazy-loaded; add scroll/wait
Anti-bot / CAPTCHAInform user; suggest running with headless: false or using their own browser session via actionbook setup extension mode

© actionbook, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/extract of actionbook/actionbook.

Open the folder on GitHubat commit 0e31254

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in actionbook/actionbook, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Extract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extract compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extract this skillactionbook/actionbook1.6k2 repos~3.4kAutomated safety check: PassApache-2.0
Cloudflare Browser Renderingeinverne/dotfiles121—~4.9kAutomated safety check: PassGPL-3.0
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone
Anti Detect Browserantibrow/anti-detect-browser-skills914—~9.8kAutomated safety check: WarnMIT
Cloudflare Browser Renderingsecondsky/claude-skills227—~4kAutomated safety check: PassMIT
Using Webctloaustegard/claude-skills150—~1.4kAutomated safety check: PassMIT

Similar skills

  • Guide for implementing Cloudflare Browser Rendering - a headless browser automation API for screenshots, PDFs, web scraping, and testing.

    121 GitHub stars~4.9k tokensUpdated 28 days ago
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    914 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings
  • Cloudflare Browser Rendering

    secondsky/claude-skills

    Cloudflare Browser Rendering with Puppeteer/Playwright. An agent skill from secondsky/claude-skills.

    227 GitHub stars~4k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Using Webctl

    oaustegard/claude-skills

    Browser automation via webctl CLI in Claude.ai containers with authenticated proxy support.

    150 GitHub stars~1.4k tokensUpdated 5 days ago
    Productivity & AutomationAuto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed

More from actionbook/actionbook

All 13 skills in this repo
  • Actionbook

    actionbook/actionbook

    Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents.

    1.6k GitHub stars~1.5k tokensUpdated 29 days ago
    Auto-check passed
  • Actionbook Web Test

    actionbook/actionbook

    Run browser-based web tests against websites using Actionbook CLI.

    1.6k GitHub stars~9.7k tokensUpdated 29 days ago
    Auto-check passed
  • Actionbook

    actionbook/actionbook

    Browser action engine. An agent skill from actionbook/actionbook.

    1.6k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Arxiv Viewer

    actionbook/actionbook

    View, search, and download academic papers from arXiv. An agent skill from actionbook/actionbook.

    1.6k GitHub stars~2k tokensUpdated 29 days ago
    Auto-check passed
  • JSON UI

    actionbook/actionbook

    CRITICAL: Use for json-ui component rendering and development.

    1.6k GitHub stars~2.7k tokensUpdated 29 days ago
    Auto-check passed
  • Active Research

    actionbook/actionbook

    Deep research and analysis tool. An agent skill from actionbook/actionbook.

    1.6k GitHub starsUsed in 1 repo~8.2k tokens
    Auto-check passed

Works with

Questions about Extract

What does Extract do?

Extract structured data from websites and produce an executable Playwright script plus extracted data. Extract is an agent skill from actionbook/actionbook. Extract structured data from websites and produce an executable Playwright script plus extracted data.

When should I use Extract?

Extract fits situations like: the user wants to scrape; harvest data from any website — product listings; any repeating content.

How do I install Extract in Claude Code?

Run `npx skills add actionbook/actionbook --skill extract -a claude-code`. Or copy the skill folder (skills/extract in actionbook/actionbook) into .claude/skills/extract in your project. Claude Code loads it when a task matches its description.

How do I install Extract in Codex?

Run `npx skills add actionbook/actionbook --skill extract -a codex`. Or copy the skill folder (skills/extract in actionbook/actionbook) into .agents/skills/extract in your project. Codex loads it when a task matches its description.

Can I use Extract in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add actionbook/actionbook --skill extract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract, .gemini/skills/extract, .github/skills/extract and .opencode/skills/extract in your project.

What does Extract need to run?

Going by SKILL.md and its folder, Extract needs the command-line tools its instructions call (node).

Does Extract access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Extract safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extract use?

Extract is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extract use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extract?

Skills that share tags, products or a category with Extract: Cloudflare Browser Rendering (einverne/dotfiles, 121 stars), Playwright Bowser (disler/bowser, 265 stars), Anti Detect Browser (antibrow/anti-detect-browser-skills, 914 stars) and Cloudflare Browser Rendering (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extract?

actionbook (a GitHub organization) maintains it in actionbook/actionbook, which has 1,609 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on September 8, 2026.

Source: actionbook/actionbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.