Data Cleaning
ericrisco/rsc-harness
A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…
Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library.
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mvanhorn/printing-press-library pp-context-dev --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .claude/skills && cp -r skills-src/library/ai/context-dev .claude/skills/pp-context-dev && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .claude/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-devType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mvanhorn/printing-press-library pp-context-dev --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .agents/skills && cp -r skills-src/library/ai/context-dev .agents/skills/pp-context-dev && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .agents/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mvanhorn/printing-press-library pp-context-dev --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/library/ai/context-dev .cursor/skills/pp-context-dev && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .cursor/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mvanhorn/printing-press-library.git --path library/ai/context-dev--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mvanhorn/printing-press-library pp-context-dev --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/library/ai/context-dev .gemini/skills/pp-context-dev && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .gemini/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mvanhorn/printing-press-library pp-context-devInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .github/skills && cp -r skills-src/library/ai/context-dev .github/skills/pp-context-dev && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .github/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mvanhorn/printing-press-library pp-context-dev --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/library/ai/context-dev .opencode/skills/pp-context-dev && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pp-context-dev" agent skill from https://github.com/mvanhorn/printing-press-library/tree/main/library/ai/context-dev into .opencode/skills/pp-context-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pp-context-dev", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pp-context-devPrinting Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library.
Pp Context Dev is an agent skill from mvanhorn/printing-press-library. Printing Press CLI for Context.dev. Agent-friendly website intelligence, brand enrichment, scraping, crawling, structured extraction, screenshots, styleguides, competitor maps, source packs, and change digests.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 146 other files (for example `.golangci.yml`, `.goreleaser.yaml` and `.manuscripts/20260623-222307/proofs/generator-version.json`).
It sits in Data & Analytics, covering Web scraping and Document parsing. The repository describes itself as: Official library of CLIs generated by the CLI Printing Press. Endorsed, tested, and community-contributed. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 7638ad4. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
goclaudenpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
CONTEXT_DEV_API_KEYCONTEXT_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pp Context Dev loads about 4.3k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,855 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mvanhorn/printing-press-library at commit 7638ad4, republished under its Apache-2.0 licence (© mvanhorn). 1,855 words, ~4,286 tokens.
.claude/skills/pp-context-dev/SKILL.md (or your agent's skills folder). This skill also uses 141 other files; get the full folder from GitHub.This skill drives the context-dev-pp-cli binary. You must verify the CLI is installed before invoking any command from this skill. If it is missing, install it first:
$HOME/.local/bin on macOS/Linux and %LOCALAPPDATA%\Programs\PrintingPress\bin on Windows:npx -y @mvanhorn/printing-press-library install context-dev --cli-onlycontext-dev-pp-cli --version$PATH for the agent/runtime that will invoke this skill.If the npx install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.6 or newer). This installs into $GOPATH/bin (default $HOME/go/bin), so add that directory to $PATH instead:
go install github.com/mvanhorn/printing-press-library/library/ai/context-dev/cmd/context-dev-pp-cli@latestIf --version reports "command not found" after install, the runtime cannot see the binary directory on $PATH. Do not proceed with skill commands until verification succeeds.
Agent-friendly Context.dev CLI for website intelligence, brand enrichment, scraping, crawling, structured extraction, screenshots, styleguides, competitor maps, source packs, and change digests.
first-class workflows — Prefer these for common Context.dev operator tasks.
context-dev-pp-cli scrape <url> --agent — Scrape one URL to clean Markdown.context-dev-pp-cli crawl <seed> --max-pages 5 --agent — Crawl same-domain pages. Use --estimate; above 25 pages requires --confirm or --yes.context-dev-pp-cli extract <url> --schema schema.json --agent — Extract typed JSON using a JSON Schema file.context-dev-pp-cli styleguide <domain|url> --agent — Extract design-system data.context-dev-pp-cli screenshot <domain|url> --agent — Capture a screenshot preview.context-dev-pp-cli entity-discover --type company --name "<name>" --location "<place>" --agent — Fielded public entity discovery with search, brand/scrape enrichment, ranking, and provenance. Valid types: company, venue, provider, school, agency, other.context-dev-pp-cli brand-brief <domain|url> --agent — Normalized brand profile with domain, website, title, description, logo, colors, fonts, socials, contact surfaces, screenshot, summary, and provenance.context-dev-pp-cli competitor-map --domain <domain|url> --max 5 --agent — Competitor set clustered by category/market with why_ranked, overlap signals, and provenance. Use --query when no seed domain exists.context-dev-pp-cli crawl-budget-plan <seed> --max-pages 25 --agent — Estimate urlRegex, max pages, credits, risks, and recommended crawl command without spending crawl credits.context-dev-pp-cli source-pack --query "<query>" --max-sources 5 --agent — Search and scrape cited source packs. Add --schema schema.json to run extraction per source. Search failure is non-zero; zero results are status: "no_results".context-dev-pp-cli website-change-digest <domain|url> --agent — Snapshot scrape/styleguide/screenshot in the CLI state dir and diff against the prior local snapshot.context-dev-pp-cli schema-lab --url <url> --url <url> --schema schema.json --agent — Run an extraction schema across samples and report field fill rates, parse failures, misses, raw statuses, and provenance.context-dev-pp-cli brand-kit <domain|url> --agent — On-demand brand kit (alias: asset-pack): logo, palette, fonts, styleguide, screenshot, favicon, socials, provenance.context-dev-pp-cli brand-qa <domain|url> --question "What is the return policy?" --agent — Ask a natural-language question about a brand's website and get a grounded answer with the URLs analyzed and provenance.context-dev-pp-cli email-enrich founders@example.com --agent — Turn a work email into a company profile and signup-form prefill fields (company name, website, industry) with provenance.context-dev-pp-cli ticker-enrich AAPL --agent — Resolve a public company from a stock ticker or ISIN to a brand profile plus NAICS/SIC industry codes. Auto-detects ticker vs ISIN; pass --exchange to disambiguate a ticker.context-dev-pp-cli trust-check <domain|url> --agent — Consistency/risk signal report across website, socials, address, phone, logo, title/domain, and web signals. Do not describe output as fraud determination.context-dev-pp-cli lead-enrich-batch leads.csv --domain-column domain --name-column name --location-column location --output enriched.json --resume --agent — Batch-enrich CSV rows with per-row success/error, provenance, and failure reason. Use --strict only when one bad row should fail the batch.For the second-wave workflows, prefer --agent for machine-stable JSON. Multi-credit commands support --estimate; all workflows support --dry-run. Estimate and dry-run responses contain estimated_credits and planned_requests and do not spend credits. website-change-digest persists snapshots under the resolved CLI state directory, never under the repo. entity-discover rejects free-form blobs and sensitive/person-identifying context; pass only public field values.
brand — Manage brand
context-dev-pp-cli brand create — Signal that you may fetch brand data for a particular domain soon to improve latency.context-dev-pp-cli brand create-ai — Given a single URL, determines if it is a product page and extracts the product information.context-dev-pp-cli brand create-ai-2 — Extract product information from a brand's website.context-dev-pp-cli brand create-ai-3 — Use AI to extract specific data points from a brand's website.context-dev-pp-cli brand create-prefetchbyemail — Signal that you may fetch brand data for a particular domain soon to improve latency.context-dev-pp-cli brand list — Retrieve logos, backdrops, colors, industry, description, and more from any domaincontext-dev-pp-cli brand list-retrievebyemail — Retrieve brand information using an email address while detecting disposable and free email addresses.context-dev-pp-cli brand list-retrievebyisin — Retrieve brand information using an ISIN (International Securities Identification Number).context-dev-pp-cli brand list-retrievebyname — Retrieve brand information using a company name.context-dev-pp-cli brand list-retrievebyticker — Retrieve brand information using a stock ticker symbol.context-dev-pp-cli brand list-retrievesimplified — Returns a simplified version of brand data containing only essential information: domain, title, colors, logoscontext-dev-pp-cli brand list-transactionidentifier — Endpoint specially designed for platforms that want to identify transaction data by the transaction title.people — Manage people
context-dev-pp-cli people — Retrieve and normalize a person profile from identifiers.web — Manage web
context-dev-pp-cli web create — Performs a crawl starting from a given URL, extracts page content as Markdowncontext-dev-pp-cli web create-extract — Crawl a website, use the provided JSON Schema and instructions to prioritize relevant internal linkscontext-dev-pp-cli web create-search — Search the web and optionally scrape each result to Markdown in one round-trip.context-dev-pp-cli web list — Analyze a company's landing page and web search evidence to return direct competitors for the same product or market.context-dev-pp-cli web list-fonts — Scrape font information from a website including font families, usage statistics, fallbacks, and element/word counts.context-dev-pp-cli web list-naics — Classify any brand into 2022 NAICS industry codes from its domain or name.context-dev-pp-cli web list-scrape — Scrapes the given URL and returns the raw HTML content of the page.context-dev-pp-cli web list-scrape-2 — Extract image assets from a web page, including standard URLs, inline SVGs, data URIs, responsive image sourcescontext-dev-pp-cli web list-scrape-3 — Scrapes the given URL into LLM usable Markdown.context-dev-pp-cli web list-scrape-4 — Crawl an entire website's sitemap and return all discovered page URLs.context-dev-pp-cli web list-screenshot — Capture a screenshot of a website.context-dev-pp-cli web list-sic — Classify any brand into Standard Industrial Classification (SIC) codes from its domain or name.context-dev-pp-cli web list-styleguide — Extract a comprehensive design system from a website including colors, typography, spacing, shadows, and UI components.When you know what you want to do but not which command does it, ask the CLI directly:
context-dev-pp-cli which "<capability in your own words>"which resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code 0 means at least one match; exit code 2 means no confident match — fall back to --help or use a narrower query.
Run context-dev-pp-cli auth setup for the URL and steps to obtain a token (add --launch to open the URL). Then store it:
context-dev-pp-cli auth set-token YOUR_TOKEN_HEREOr set CONTEXT_DEV_API_KEY as an environment variable. CONTEXT_API_KEY is accepted as a fallback when CONTEXT_DEV_API_KEY is unset; the legacy generated CONTEXT_DEV_BEARER_AUTH variable remains supported for compatibility.
Run context-dev-pp-cli doctor to verify setup.
Add --agent to any command. Expands to: --json --compact --no-input --no-color --yes.
Pipeable — JSON on stdout, errors on stderr
Filterable — --select keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:
context-dev-pp-cli brand list --domain example-value --agent --select id,name,statusPreviewable — --dry-run shows the request without sending
Offline-friendly — sync/search commands can use the local SQLite store when available
Non-interactive — never prompts, every input is a flag
Explicit retries — use --idempotent only when an already-existing create should count as success
Commands that read from the local store or the API wrap output in a provenance envelope:
{
"meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."},
"results": <data>
}Parse .results for data and .meta.source to know whether it's live or local. A human-readable N results (live) summary is printed to stderr only when stdout is a terminal AND no machine-format flag (--json, --csv, --compact, --quiet, --plain, --select) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout.
Agents should treat the CLI's path resolver as part of the runtime contract:
Use --home <dir> for one invocation, or set CONTEXT_DEV_HOME=<dir> to relocate all four path kinds under one root.
Use per-kind env vars only when a specific kind must diverge: CONTEXT_DEV_CONFIG_DIR, CONTEXT_DEV_DATA_DIR, CONTEXT_DEV_STATE_DIR, CONTEXT_DEV_CACHE_DIR.
Resolution order is per-kind env var, --home, CONTEXT_DEV_HOME, XDG (XDG_CONFIG_HOME, XDG_DATA_HOME, XDG_STATE_HOME, XDG_CACHE_HOME), then platform defaults.
config contains settings like config.toml and profiles. data contains credentials.toml, data.db, cookies, and auth sidecars. state contains persisted queries, jobs, and teach.log. cache contains regenerable HTTP/cache files.
Stored secrets live in credentials.toml under the data dir. Existing legacy config.toml secrets are read for compatibility and leave config.toml on the first auth write.
Run context-dev-pp-cli doctor --fail-on warn to surface path and credential-location warnings. agent-context exposes a schema v4 paths block for agents that need the resolved dirs.
For MCP, pass relocation through the MCP host config. The MCP binary does not inherit CLI flags:
{
"mcpServers": {
"context-dev": {
"command": "context-dev-pp-mcp",
"env": {
"CONTEXT_DEV_HOME": "/srv/context-dev"
}
}
}
}Fleet precedence: an inherited per-kind env var overrides an explicit --home for that kind. Use CONTEXT_DEV_HOME or per-kind vars as durable fleet levers, and use --home only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing CONTEXT_DEV_HOME, or doctor will not find credentials left under the former root.
When you (or the agent) notice something off about this CLI, record it:
context-dev-pp-cli feedback "the --since flag is inclusive but docs say exclusive"
context-dev-pp-cli feedback --stdin < notes.txt
context-dev-pp-cli feedback list --json --limit 10Entries are stored locally as feedback.jsonl under the resolved data dir. They are never POSTed unless CONTEXT_DEV_FEEDBACK_ENDPOINT is set AND either --send is passed or CONTEXT_DEV_FEEDBACK_AUTO_SEND=true. Default behavior is local-only.
Write what surprised you, not a bug report. Short, specific, one line: that is the part that compounds.
Every command accepts --deliver <sink>. The output goes to the named sink in addition to (or instead of) stdout, so agents can route command results without hand-piping. Three sinks are supported:
| Sink | Effect |
|---|---|
stdout | Default; write to stdout only |
file:<path> | Atomically write output to <path> (tmp + rename) |
webhook:<url> | POST the output body to the URL (application/json or application/x-ndjson when --compact) |
Unknown schemes are refused with a structured error naming the supported set. Webhook failures return non-zero and log the URL + HTTP status on stderr.
A profile is a saved set of flag values, reused across invocations. Use it when a scheduled agent calls the same command every run with the same configuration - HeyGen's "Beacon" pattern.
context-dev-pp-cli profile save briefing --json
context-dev-pp-cli --profile briefing brand list --domain example-value
context-dev-pp-cli profile list --json
context-dev-pp-cli profile show briefing
context-dev-pp-cli profile delete briefing --yesExplicit flags always win over profile values; profile values win over defaults. agent-context lists all available profiles under available_profiles so introspecting agents discover them at runtime.
| Code | Meaning |
|---|---|
| 0 | Success |
| 2 | Usage error (wrong arguments) |
| 3 | Resource not found |
| 4 | Authentication required |
| 5 | API error (upstream issue) |
| 7 | Rate limited (wait and retry) |
| 10 | Config error |
Parse $ARGUMENTS:
help, or --help → show context-dev-pp-cli --help outputinstall → ends with mcp → MCP installation; otherwise → see Prerequisites above--agent)go install github.com/mvanhorn/printing-press-library/library/ai/context-dev/cmd/context-dev-pp-mcp@latestclaude mcp add context-dev-pp-mcp -- context-dev-pp-mcpclaude mcp listwhich context-dev-pp-cli
If not found, offer to install (see Prerequisites at the top of this skill).--agent flag:context-dev-pp-cli <command> [subcommand] [args] --agentcontext-dev-pp-cli <command> --help.© mvanhorn, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 141 other files in library/ai/context-dev of mvanhorn/printing-press-library.
Open the folder on GitHubat commit 7638ad4
Pp Context Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pp Context Dev this skillmvanhorn/printing-press-library | 2.1k | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | |
| Data Cleaningericrisco/rsc-harness | 156 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Data Scraperericrisco/rsc-harness | 156 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Cheerio ParsingKilo-Org/kilo-marketplace | 189 | 1 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Steel Browseraiskillstore/marketplace | 430 | — | ~2k | Automated safety check: Pass | None | |
| Optimize Button Patternsduckduckgo/tracker-radar-collector | 169 | — | ~1.6k | Automated safety check: Pass | Custom licence |
ericrisco/rsc-harness
A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…
ericrisco/rsc-harness
A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…
Kilo-Org/kilo-marketplace
Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.
aiskillstore/marketplace
Use this skill by default for browser or web tasks that can run in the cloud: site navigation, scraping, structured extraction, screenshots/PDFs, form flows, and anti-bot-sensitive automation.
duckduckgo/tracker-radar-collector
Iteratively improves cookie-popup button regex patterns in button-patterns.js against labelled-button-texts.csv.
zhaojiaqi/MeowHub
Browse the web using Browserless.io cloud browser service. An agent skill from zhaojiaqi/MeowHub.
mvanhorn/printing-press-library
Desktop automation through the real Rust agent-desktop CLI, published in Printing Press through a small bridge.
mvanhorn/printing-press-library
Search, browse, and download Google Fonts from the terminal via the gfonts CLI.
mvanhorn/printing-press-library
The free, offline Trigger phrases: search 1688 for, find a factory on 1688 for, wholesale price on 1688 for, who is the cheapest supplier on 1688 for, compare 1688 suppliers for, use 1688, run 1688.
mvanhorn/printing-press-library
Inspect known Activity Japan plan IDs or URLs, compare dated prices and sessions, check language-sitemap coverage, and hand off to canonical booking pages.
mvanhorn/printing-press-library
Every Admin By Request portal action, plus a local SQLite mirror of audit, events, inventory and requests for ad-hoc...
mvanhorn/printing-press-library
macOS screen capture, window recording, GIF conversion, and agent evidence bundles from the terminal.
Categories
Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library. Pp Context Dev is an agent skill from mvanhorn/printing-press-library.dev.
Pp Context Dev fits situations like: tasks that involve Web scraping; tasks that involve Document parsing.
Run `npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a claude-code`. Or copy the skill folder (library/ai/context-dev in mvanhorn/printing-press-library) into .claude/skills/pp-context-dev in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a codex`. Or copy the skill folder (library/ai/context-dev in mvanhorn/printing-press-library) into .agents/skills/pp-context-dev in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mvanhorn/printing-press-library --skill pp-context-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pp-context-dev, .gemini/skills/pp-context-dev, .github/skills/pp-context-dev and .opencode/skills/pp-context-dev in your project.
Going by SKILL.md and its folder, Pp Context Dev needs the command-line tools its instructions call (go, claude and npx) and credentials named CONTEXT_DEV_API_KEY and CONTEXT_API_KEY. Our summary lists: Node.js; A credential in CONTEXT_DEV_API_KEY; A credential in CONTEXT_API_KEY. Its frontmatter pre-approves these tools: Read, Bash.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Pp Context Dev is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pp Context Dev: Data Cleaning (ericrisco/rsc-harness, 156 stars), Data Scraper (ericrisco/rsc-harness, 156 stars), Cheerio Parsing (Kilo-Org/kilo-marketplace, 189 stars) and Steel Browser (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mvanhorn (a GitHub user) maintains it in mvanhorn/printing-press-library, which has 2,053 GitHub stars. The repository holds 505 skills in this directory. The repository was last updated on October 6, 2026.
Source: mvanhorn/printing-press-library on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.