Scraper Builder
jwynia/agent-skills
Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment.
Extracts icons, metadata, text, forms, videos and social links from any web page with playwright-cli, with SVG icon classification and cleanup.
$ npx skills add adobe/skills --skill page-collect -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install adobe/skills page-collect --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/web/skills/page-collect .claude/skills/page-collect && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .claude/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collectType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add adobe/skills --skill page-collect -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install adobe/skills page-collect --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/web/skills/page-collect .agents/skills/page-collect && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .agents/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adobe/skills --skill page-collect -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install adobe/skills page-collect --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/web/skills/page-collect .cursor/skills/page-collect && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .cursor/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/adobe/skills.git --path plugins/web/skills/page-collect--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add adobe/skills --skill page-collect -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install adobe/skills page-collect --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/web/skills/page-collect .gemini/skills/page-collect && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .gemini/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install adobe/skills page-collectInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add adobe/skills --skill page-collect -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/web/skills/page-collect .github/skills/page-collect && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .github/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adobe/skills --skill page-collect -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install adobe/skills page-collect --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/web/skills/page-collect .opencode/skills/page-collect && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "page-collect" agent skill from https://github.com/adobe/skills/tree/main/plugins/web/skills/page-collect into .opencode/skills/page-collect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "page-collect", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
page-collectExtracts icons, metadata, text, forms, videos and social links from any web page with playwright-cli, with SVG icon classification and cleanup.
This skill extracts structured resources from any web page with `playwright-cli`: icons, metadata, body text, forms, videos and social links. A bundled `page-collect.js` script takes a subcommand (`all`, `icons`, `metadata`, `text`, `forms`, `videos` or `socials`) and a URL, and writes JSON files such as `metadata.json` and `forms.json`, an `icons/` folder and, for `all`, a `collection.json` and a screenshot, into `page-collect-output/` unless you choose another directory.
The icon collector is the most detailed part. It finds SVGs in inline elements, `img` tags, CSS background data URIs and sprite references, then classifies each as an icon (small and inside a button, link or nav), a logo (in a brand area) or an image (larger and standalone, which is excluded). Names come from DOM context such as aria labels, classes and IDs, and unnamed ones get numbered names flagged with low name confidence for you to review. Each icon SVG is cleaned: declarations, comments and metadata stripped, a viewBox ensured, fixed sizes removed, and fills and strokes switched to currentColor, except for logos.
It needs Node 22 or newer and `playwright-cli` on the PATH, and it can take a browser recipe from the `browser-probe` skill for pages behind bot protection. Icons are optimized for EDS and land in `/icons/` for use with `decorateIcons()`.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4c67484. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Node 22+ and playwright-cli on PATH. Run `playwright-cli --help` for usage.
From compatibility in the SKILL.md frontmatter.
Web Page Resource Collector loads about 1k tokens when it runs, and up to ~2.8k if it reads all its reference files. Until then it costs about 118 tokens; SKILL.md has 360 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from adobe/skills at commit 4c67484, republished under its Apache-2.0 licence (© adobe). 360 words, ~1,045 tokens.
.claude/skills/page-collect/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Extract structured resources from any webpage via playwright-cli.
Node 22+ required. Run playwright-cli --help for the command reference.
| Subcommand | Purpose | Output |
|---|---|---|
all | Run all collectors | collection.json, screenshot.jpg + assets |
icons | SVGs, icon fonts, CSS icons → classified SVGs | icons/ + icons.json |
metadata | Meta tags, OG, structured data | metadata.json |
text | Body text, headings, word count | text.json |
forms | Form structures, fields, actions | forms.json |
videos | Video embeds, sources | videos.json |
socials | Social media links | socials.json |
If CLAUDE_SKILL_DIR is set:
SCRIPT="${CLAUDE_SKILL_DIR}/scripts/page-collect.js"Otherwise, find it:
SCRIPT="$(find ~/.claude -path "*/page-collect/scripts/page-collect.js" -type f 2>/dev/null | head -1)"node "$SCRIPT" <subcommand> <url> [--output <dir>]Default output: ./page-collect-output/
playwright-cli must be on PATH. Optionally pass --browser-recipe <path> to
use a browser-recipe.json from the browser-probe skill to bypass bot protection.
The icon collector extracts SVGs from multiple sources:
<svg> elements<img> tags with .svg src or data:image/svg+xml URIsbackground-image SVG data URIs<use> sprite references (resolved to standalone SVGs)| Class | Criteria | Output |
|---|---|---|
icon | ≤ 48px, inside button/link/nav | /icons/{name}.svg |
logo | Brand area, "logo" in class/alt/src | /icons/logo.svg |
image | > 48px, standalone | Excluded |
Icons are named from DOM context (aria-label, class, ID). When no
meaningful name can be derived, they get icon-{n} with
nameConfidence: "low" in the manifest — review these and rename.
Each icon SVG is cleaned:
currentColor (icons only, not logos)For more details, read the collectors reference in references/collectors.md.
{
"url": "https://example.com",
"icons": [
{
"name": "search",
"class": "icon",
"source": "inline-svg",
"file": "icons/search.svg",
"nameConfidence": "high",
"context": "header button Search"
}
]
}icons.json — rename any nameConfidence: "low" icons/icons/*.svg to the EDS project's /icons/ directory:iconname: notationdecorateIcons() in aem.js handles renderingall results:Review collection.json for a full resource inventory of the page.
When used as part of a header migration:
node "$SCRIPT" icons <source-url> --output <extraction-dir>icons.json and copies SVGs to /icons/nav.plain.html uses :iconname: for tools/utility iconsprogram.md notes available icons© adobe, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in plugins/web/skills/page-collect of adobe/skills.
Open the folder on GitHubat commit 4c67484
Web Page Resource Collector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Page Resource Collector this skilladobe/skills | 197 | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Scraper Builderjwynia/agent-skills | 169 | — | ~4k | Automated safety check: Pass | MIT | |
| Skyvern Browser AutomationSkyvern-AI/skyvern | 23k | 1 repos | ~2.9k | Automated safety check: Pass | AGPL-3.0 | |
| Camofox Browserredf0x1/camofox-browser | 412 | — | ~4.6k | Automated safety check: Pass | MIT | |
| Playwright Bowserdisler/bowser | 265 | — | ~1.1k | Automated safety check: Notes | None | |
| Web Crawlerbyungjunjang/web-crawler | 165 | — | ~7.2k | Automated safety check: Pass | MIT |
jwynia/agent-skills
Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment.
Skyvern-AI/skyvern
Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching.
redf0x1/camofox-browser
Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.
disler/bowser
Headless browser automation using Playwright CLI. An agent skill from disler/bowser.
byungjunjang/web-crawler
URL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트. An agent skill from byungjunjang/web-crawler.
code-yeongyu/oh-my-openagent
Routes a web request through the cheapest tier that can finish it, from headless extraction with WAF bypass up to a real stealth or signed-in browser, with screenshots as proof.
adobe/skills
Scaffolds, implements, deploys and debugs Adobe Runtime actions in App Builder projects, with templates for webhooks, events, database CRUD, sequences and Asset Compute workers.
adobe/skills
Launches Chrome with an unpacked extension over CDP, opens its sidepanel, popup or options page, and hands over to cdp-connect for clicks, typing and screenshots.
adobe/skills
Detect all languages used on a webpage — both declared (html@lang, hreflang alternate links, nested lang= attributes, meta content-language) and actually present in the body text (Google CLD3 via…
adobe/skills
Prepare any webpage for clean interaction by detecting and removing disruptive overlays (cookie banners, GDPR consent, modals, popups, newsletter signups, paywalls, login walls).
adobe/skills
Reduce a webpage to a structural skeleton with semantic tokens.
adobe/skills
Use this when converting an AI-generated static HTML page (Stardust, Mobirise, Relume, Lovable, v0, Figma-derived, etc.) into an Edge Delivery Services page while preserving the original design and…
Works with
Categories
Extracts icons, metadata, text, forms, videos and social links from any web page with playwright-cli, with SVG icon classification and cleanup. This skill extracts structured resources from any web page with `playwright-cli`: icons, metadata, body text, forms, videos and social links.json` and a screenshot, into `page-collect-output/` unless you choose another directory.
Web Page Resource Collector fits situations like: migrating a page and needing its icons, text and metadata as files; auditing a site's forms, videos or social links; pulling a page's SVG icons and logo into an icons folder.
Run `npx skills add adobe/skills --skill page-collect -a claude-code`. Or copy the skill folder (plugins/web/skills/page-collect in adobe/skills) into .claude/skills/page-collect in your project. Claude Code loads it when a task matches its description.
Run `npx skills add adobe/skills --skill page-collect -a codex`. Or copy the skill folder (plugins/web/skills/page-collect in adobe/skills) into .agents/skills/page-collect in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adobe/skills --skill page-collect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/page-collect, .gemini/skills/page-collect, .github/skills/page-collect and .opencode/skills/page-collect in your project.
Going by SKILL.md and its folder, Web Page Resource Collector needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Node 22 or newer; `playwright-cli` on the PATH. Compatibility (from SKILL.md): Requires Node 22+ and playwright-cli on PATH. Run `playwright-cli --help` for usage..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Web Page Resource Collector is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Web Page Resource Collector: Scraper Builder (jwynia/agent-skills, 169 stars), Skyvern Browser Automation (Skyvern-AI/skyvern, 23k stars), Camofox Browser (redf0x1/camofox-browser, 412 stars) and Playwright Bowser (disler/bowser, 265 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
adobe (a GitHub organization) maintains it in adobe/skills, which has 197 GitHub stars. The repository holds 65 skills in this directory. The repository was last updated on October 9, 2026.
Source: adobe/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.