Agent skill

Pp Scrape Do

by mvanhorn in mvanhorn/printing-press-library

The first CLI for Scrape-do with Google SERP scraping plus a credit and concurrency governor Trigger phrases: scrape google for, get google search results for, track keyword rank for, scrape this…

Apache-2.0Auto-check: notesData & Analytics

Install Pp Scrape Do

skills CLI
$ npx skills add mvanhorn/printing-press-library --skill pp-scrape-do -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mvanhorn/printing-press-library pp-scrape-do --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-skills/pp-scrape-do .claude/skills/pp-scrape-do && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pp-scrape-do
GitHub stars
2.1k
Token cost
~3.2k tokens
SKILL.md length
1,400 words
Files
1
Skills in repo
505
Repo updated
First seen
Licence
Apache-2.0

At a glance

The first CLI for Scrape-do with Google SERP scraping plus a credit and concurrency governor Trigger phrases: scrape google for, get google search results for, track keyword rank for, scrape this…

  • Works in 3 steps: Install via the Printing Press installer → Verify: scrape-do-pp-cli --version → Ensure the directory the npx installer…
  • Phrases: scrape google for
  • SKILL.md covers Prerequisites: Install the CLI, When to Use This CLI, When Not to Use This CLI and Unique Capabilities, plus 11 more sections
  • Calls go, claude and npx; reaches linkedin.com; needs SCRAPEDO_API_KEY

What it does

Pp Scrape Do is an agent skill from mvanhorn/printing-press-library. The first CLI for Scrape-do with Google SERP scraping plus a credit and concurrency governor Trigger phrases: scrape google for, get google search results for, track keyword rank for, scrape this url with scrape.do, estimate scrape.do cost for, check my scrape.do credits, use scrape-do, run scrape-do.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping and Web search. The repository describes itself as: Official library of CLIs generated by the CLI Printing Press. Endorsed, tested, and community-contributed. The licence is Apache-2.0.

When your agent uses it

  • Phrases: scrape google for
  • Get google search results for
  • Track keyword rank for
  • Scrape this url with scrape.do

Example prompts

  • “/pp-scrape-do”

Requirements

  • Node.js
  • A credential in SCRAPEDO_API_KEY
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Install via the Printing Press installer
  2. Verify: scrape-do-pp-cli --version
  3. Ensure the directory the npx installer placed the binary in is on $PATH. On Windows this is %LOCALAPPDATA%\Programs\PrintingPress\bin; on…

What it can do on your machine

Read from SKILL.md and the folder at commit 0fdcc7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go
    • claude
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • linkedin.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SCRAPEDO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pp Scrape Do loads about 3.2k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,400 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mvanhorn/printing-press-library at commit 0fdcc7a, republished under its Apache-2.0 licence (© mvanhorn). 1,400 words, ~3,234 tokens.

Download SKILL.mdSave it as .claude/skills/pp-scrape-do/SKILL.md (or your agent's skills folder).
name
pp-scrape-do
description
The first CLI for Scrape-do with Google SERP scraping plus a credit and concurrency governor Trigger phrases: `scrape google for`, `get google search results for`, `track keyword rank for`, `scrape this url with scrape.do`, `estimate scrape.do cost for`, `check my scrape.do credits`, `use scrape-do`, `run scrape-do`.
allowed-tools
Read, Bash
author
Charles Garrison
license
Apache-2.0
argument-hint
<command> [args] | install cli|mcp
<!-- GENERATED FILE — DO NOT EDIT.
     This file is a verbatim mirror of library/developer-tools/scrape-do/SKILL.md,
     regenerated post-merge by tools/generate-skills/. Hand-edits here are
     silently overwritten on the next regen. Edit the library/ source instead.
     See the repository agent guide, section "Generated artifacts: registry.json, cli-skills/". -->

Scrape.do — Printing Press CLI

Prerequisites: Install the CLI

This skill drives the scrape-do-pp-cli binary. You must verify the CLI is installed before invoking any command from this skill. If it is missing, install it first:

  1. Install via the Printing Press installer:
    bash
    npx -y @mvanhorn/printing-press-library install scrape-do --cli-only
  2. Verify: scrape-do-pp-cli --version
  3. Ensure the directory the npx installer placed the binary in is on $PATH. On Windows this is %LOCALAPPDATA%\Programs\PrintingPress\bin; on macOS/Linux it is typically $HOME/.local/bin or a platform-specific location printed by the installer. If you used the Go fallback instead, add $GOPATH/bin (or $HOME/go/bin).

If the npx install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.6 or newer):

bash
go install github.com/mvanhorn/printing-press-library/library/developer-tools/scrape-do/cmd/scrape-do-pp-cli@latest

If --version reports "command not found" after install, the install step did not put the binary on $PATH. Do not proceed with skill commands until verification succeeds.

Every Scrape.do surface — the core scraper plus the whole Google family (search, maps, news, shopping, flights, hotels, trends) — wrapped with an offline SQLite history. The governor estimates credit cost before every call with cost, debits a local ledger from the authoritative cost header with budget, and gates concurrent requests against your plan's live ceiling so an agent swarm never 429s itself or burns the monthly budget. SERPs become queryable history: drift and movers surface rank changes offline with no re-spend.

When to Use This CLI

Reach for this CLI whenever a task needs Google search results, structured SERP data, or any proxied web scrape via Scrape.do — and especially when several agents share one Scrape.do account. Its governor (cost/budget/batch) keeps concurrent usage inside the plan's limits and the monthly credit budget, and its offline SERP history (drift/movers) answers rank-change questions without re-spending credits. Prefer it over raw curl calls when you care about credit cost, concurrency safety, or comparing results over time.

When Not to Use This CLI

Do not activate this CLI for requests that require creating, updating, deleting, publishing, commenting, upvoting, inviting, ordering, sending messages, booking, purchasing, or changing remote state. No command mutates remote target state. Note, however, that scrape, google, and batch make billed GET requests that consume Scrape.do account credits — use cost to estimate before spending and budget to track and cap spend.

Unique Capabilities

These capabilities aren't available in any other tool for this API.

Offline SERP intelligence
  • drift — Diff a Google query's two most recent stored SERPs and see exactly which results moved, appeared, or dropped — entirely offline, no credits spent.

    When an agent needs to know whether a ranking changed since last check, this answers it without re-spending a 10-credit SERP call.

    bash
    scrape-do-pp-cli drift "best crm software" --json
  • movers — Scan every tracked query's latest-versus-previous SERP snapshot and surface only the queries whose top positions moved past a threshold.

    Turns hundreds of tracked keywords into a short 'what changed this week' list an agent can act on.

    bash
    scrape-do-pp-cli movers --threshold 3 --agent
Credit & concurrency governance
  • batch — Dispatch a list of URLs or queries through a shared concurrency lease and per-call credit ledger, auto-retrying only the non-billed 429/502/510 classes and stopping before a credit ceiling.

    Lets an agent swarm hammer one account at full speed without 429 storms or blowing the monthly budget.

    bash
    scrape-do-pp-cli batch --input urls.txt --max-credits 500 --agent
  • budget — Attribute spend by mode and query-family from a local ledger debited off the authoritative per-call cost header, joined with cached account state to forecast burn-rate against days remaining.

    Tells an agent how much budget is left and which workloads are eating it before the account hits a hard 401.

    bash
    scrape-do-pp-cli budget --agent
  • cost — Print the exact credit cost a request will incur — accounting for render, super-proxy, Google endpoints, and per-domain overrides — before spending a single credit.

    Lets an agent compare the cost of cheap vs expensive scrape modes and pick the cheapest path that works.

    bash
    scrape-do-pp-cli cost --url https://www.linkedin.com/company/example --render --super

Command Reference

account — Live Scrape.do account state: subscription status, concurrency allowance and headroom, monthly credit cap and remaining credits.

  • scrape-do-pp-cli account — Fetch live account state: IsActive, ConcurrentRequest, RemainingConcurrentRequest, MaxMonthlyRequest
Finding the right command

When you know what you want to do but not which command does it, ask the CLI directly:

bash
scrape-do-pp-cli which "<capability in your own words>"

which resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code 0 means at least one match; exit code 2 means no confident match — fall back to --help or use a narrower query.

Recipes

Narrow a deep SERP payload for an agent
bash
scrape-do-pp-cli google search "coffee makers" --agent --select organic_results.position,organic_results.title,organic_results.link

The SERP JSON is large and deeply nested; --select with dotted paths returns just rank, title, and link so an agent doesn't burn context parsing the full payload.

Estimate cost before an expensive scrape
bash
scrape-do-pp-cli cost --url https://www.linkedin.com/company/example --render --super

Prints the expected credits (LinkedIn domain override + render + super proxy) with no API spend, so you can choose the cheapest mode that works.

Fan out a URL list under the concurrency cap
bash
scrape-do-pp-cli batch --input urls.txt --max-credits 500 --agent

Dispatches every URL through the shared concurrency lease, retries only the non-billed 429/502/510 classes, and stops before the 500-credit ceiling.

See what ranked-changes happened this week
bash
scrape-do-pp-cli movers --threshold 3 --agent

Compares each tracked query's two latest stored SERPs and lists only the queries whose top positions moved by 3 or more — entirely offline.

Show full SKILL.md (598 more words)Show less
Query your stored SERP history with SQL
bash
scrape-do-pp-cli sql "SELECT domain, COUNT(*) c FROM serp_organic GROUP BY domain ORDER BY c DESC LIMIT 10"

Read-only SQL over the local store — share-of-voice by domain across every stored SERP, with no credit spent.

Auth Setup

Scrape.do uses a single API token passed as the token query parameter. Set it as SCRAPEDO_API_KEY in your environment and the CLI never logs the value. The same token drives the core scraper and every Google and Ready-API endpoint.

Run scrape-do-pp-cli doctor to verify setup.

Agent Mode

Add --agent to any command. Expands to: --json --compact --no-input --no-color --yes.

  • Pipeable — JSON on stdout, errors on stderr

  • Filterable — --select keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:

    bash
    scrape-do-pp-cli account --agent --select id,name,status
  • Previewable — --dry-run shows the request without sending

  • Offline-friendly — sync/search commands can use the local SQLite store when available

  • Non-interactive — never prompts, every input is a flag

  • Read-only — do not use this CLI for create, update, delete, publish, comment, upvote, invite, order, send, or other mutating requests

Response envelope

Commands that read from the local store or the API wrap output in a provenance envelope:

json
{
  "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."},
  "results": <data>
}

Parse .results for data and .meta.source to know whether it's live or local. A human-readable N results (live) summary is printed to stderr only when stdout is a terminal AND no machine-format flag (--json, --csv, --compact, --quiet, --plain, --select) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout.

Agent Feedback

When you (or the agent) notice something off about this CLI, record it:

scrape-do-pp-cli feedback "the --since flag is inclusive but docs say exclusive"
scrape-do-pp-cli feedback --stdin < notes.txt
scrape-do-pp-cli feedback list --json --limit 10

Entries are stored locally at ~/.local/share/scrape-do-pp-cli/feedback.jsonl. They are never POSTed unless SCRAPE_DO_FEEDBACK_ENDPOINT is set AND either --send is passed or SCRAPE_DO_FEEDBACK_AUTO_SEND=true. Default behavior is local-only.

Write what surprised you, not a bug report. Short, specific, one line: that is the part that compounds.

Output Delivery

Every command accepts --deliver <sink>. The output goes to the named sink in addition to (or instead of) stdout, so agents can route command results without hand-piping. Three sinks are supported:

SinkEffect
stdoutDefault; write to stdout only
file:<path>Atomically write output to <path> (tmp + rename)
webhook:<url>POST the output body to the URL (application/json or application/x-ndjson when --compact)

Unknown schemes are refused with a structured error naming the supported set. Webhook failures return non-zero and log the URL + HTTP status on stderr.

Named Profiles

A profile is a saved set of flag values, reused across invocations. Use it when a scheduled agent calls the same command every run with the same configuration - HeyGen's "Beacon" pattern.

scrape-do-pp-cli profile save briefing --json
scrape-do-pp-cli --profile briefing account
scrape-do-pp-cli profile list --json
scrape-do-pp-cli profile show briefing
scrape-do-pp-cli profile delete briefing --yes

Explicit flags always win over profile values; profile values win over defaults. agent-context lists all available profiles under available_profiles so introspecting agents discover them at runtime.

Exit Codes

CodeMeaning
0Success
2Usage error (wrong arguments)
3Resource not found
4Authentication required
5API error (upstream issue)
7Rate limited (wait and retry)
10Config error

Argument Parsing

Parse $ARGUMENTS:

  1. Empty, help, or --help → show scrape-do-pp-cli --help output
  2. Starts with install → ends with mcp → MCP installation; otherwise → see Prerequisites above
  3. Anything else → Direct Use (execute as CLI command with --agent)

MCP Server Installation

  1. Install the MCP server:
    bash
    go install github.com/mvanhorn/printing-press-library/library/developer-tools/scrape-do/cmd/scrape-do-pp-mcp@latest
  2. Register with Claude Code:
    bash
    claude mcp add scrape-do-pp-mcp -- scrape-do-pp-mcp
  3. Verify: claude mcp list

Direct Use

  1. Check if installed: which scrape-do-pp-cli If not found, offer to install (see Prerequisites at the top of this skill).
  2. Match the user query to the best command from the Unique Capabilities and Command Reference above.
  3. Execute with the --agent flag:
    bash
    scrape-do-pp-cli <command> [subcommand] [args] --agent
  4. If ambiguous, drill into subcommand help: scrape-do-pp-cli <command> --help.

© mvanhorn, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-skills/pp-scrape-do of mvanhorn/printing-press-library.

Open the folder on GitHubat commit 0fdcc7a

Compare with similar skills

Pp Scrape Do next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pp Scrape Do compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pp Scrape Do this skillmvanhorn/printing-press-library2.1k—~3.2kAutomated safety check: NotesApache-2.0
Agent Readiness Auditindranilbanerjee/digital-marketing-pro8551 repos~3.9kAutomated safety check: PassMIT
Web Search Scraper API Skillbrowser-act/skills6.1k1 repos~1.3kAutomated safety check: PassMIT
Keirouter Web Fetchmydisha/keirouter147—~741Automated safety check: PassMIT
Deepapidavidondrej/skills4.1k—~2.5kAutomated safety check: PassMIT
Bright Data Best Practicesbrightdata/skills2641 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended)…

    855 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • This skill helps users automatically extract complete Markdown content from any website via the BrowserAct Web Search Scraper API.

    6.1k GitHub starsUsed in 1 repo~1.3k tokens
    Data & AnalyticsAuto-check passed
  • Keirouter Web Fetch

    mydisha/keirouter

    Fetch URL → markdown / text / HTML via KeiRouter /v1/web/fetch using Firecrawl / Jina Reader / Tavily Extract / Exa Contents.

    147 GitHub stars~741 tokensUpdated 27 days ago
    Data & AnalyticsAuto-check passed
  • Deepapi

    davidondrej/skills

    Use DeepAPI for all web search, deep research, and web scraping (websites, LinkedIn, GitHub, X/Twitter, YouTube, Instagram) instead of built-in search, research, fetch, or browser tools.

    4.1k GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Build production-ready Bright Data integrations with best practices baked in.

    264 GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Search

    brightdata/skills

    Search the web via the Bright Data CLI — bdata search for Google/Bing/Yandex SERP, bdata discover for intent-ranked semantic results.

    264 GitHub stars~1.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from mvanhorn/printing-press-library

All 505 skills in this repo
  • Agent Desktop

    mvanhorn/printing-press-library

    Desktop automation through the real Rust agent-desktop CLI, published in Printing Press through a small bridge.

    2.1k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check: notes
  • Gfonts

    mvanhorn/printing-press-library

    Search, browse, and download Google Fonts from the terminal via the gfonts CLI.

    2.1k GitHub stars~574 tokensUpdated yesterday
    Auto-check passed
  • Pp 1688

    mvanhorn/printing-press-library

    The free, offline Trigger phrases: search 1688 for, find a factory on 1688 for, wholesale price on 1688 for, who is the cheapest supplier on 1688 for, compare 1688 suppliers for, use 1688, run 1688.

    2.1k GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • Pp Activity Japan

    mvanhorn/printing-press-library

    Inspect known Activity Japan plan IDs or URLs, compare dated prices and sessions, check language-sitemap coverage, and hand off to canonical booking pages.

    2.1k GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Pp Adminbyrequest

    mvanhorn/printing-press-library

    Every Admin By Request portal action, plus a local SQLite mirror of audit, events, inventory and requests for ad-hoc...

    2.1k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check: notes
  • Pp Agent Capture

    mvanhorn/printing-press-library

    macOS screen capture, window recording, GIF conversion, and agent evidence bundles from the terminal.

    2.1k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check: notes

Questions about Pp Scrape Do

What does Pp Scrape Do do?

The first CLI for Scrape-do with Google SERP scraping plus a credit and concurrency governor Trigger phrases: scrape google for, get google search results for, track keyword rank for, scrape this…. Pp Scrape Do is an agent skill from mvanhorn/printing-press-library.do credits, use scrape-do, run scrape-do.

When should I use Pp Scrape Do?

Pp Scrape Do fits situations like: phrases: scrape google for; get google search results for; track keyword rank for; scrape this url with scrape.do.

How do I install Pp Scrape Do in Claude Code?

Run `npx skills add mvanhorn/printing-press-library --skill pp-scrape-do -a claude-code`. Or copy the skill folder (cli-skills/pp-scrape-do in mvanhorn/printing-press-library) into .claude/skills/pp-scrape-do in your project. Claude Code loads it when a task matches its description.

How do I install Pp Scrape Do in Codex?

Run `npx skills add mvanhorn/printing-press-library --skill pp-scrape-do -a codex`. Or copy the skill folder (cli-skills/pp-scrape-do in mvanhorn/printing-press-library) into .agents/skills/pp-scrape-do in your project. Codex loads it when a task matches its description.

Can I use Pp Scrape Do in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mvanhorn/printing-press-library --skill pp-scrape-do -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pp-scrape-do, .gemini/skills/pp-scrape-do, .github/skills/pp-scrape-do and .opencode/skills/pp-scrape-do in your project.

What does Pp Scrape Do need to run?

Going by SKILL.md and its folder, Pp Scrape Do needs the command-line tools its instructions call (go, claude and npx) and credentials named SCRAPEDO_API_KEY. Our summary lists: Node.js; A credential in SCRAPEDO_API_KEY. Its frontmatter pre-approves these tools: Read, Bash.

Does Pp Scrape Do access the network?

SKILL.md names 1 domain. In commands or code: linkedin.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Pp Scrape Do safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Pp Scrape Do use?

Pp Scrape Do is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pp Scrape Do use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pp Scrape Do?

Skills that share tags, products or a category with Pp Scrape Do: Agent Readiness Audit (indranilbanerjee/digital-marketing-pro, 855 stars), Web Search Scraper API Skill (browser-act/skills, 6.1k stars), Keirouter Web Fetch (mydisha/keirouter, 147 stars) and Deepapi (davidondrej/skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pp Scrape Do?

mvanhorn (a GitHub user) maintains it in mvanhorn/printing-press-library, which has 2,053 GitHub stars. The repository holds 505 skills in this directory. The repository was last updated on October 7, 2026.

Source: mvanhorn/printing-press-library on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.