Agent skill

Extract Page Data

by lintendo in lintendo/Axhub-Skills

Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs.

MITAuto-check: warningsFrontend & Design

Install Extract Page Data

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add lintendo/Axhub-Skills --skill extract-page-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lintendo/Axhub-Skills extract-page-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lintendo/Axhub-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/extract-page-data .claude/skills/extract-page-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extract-page-data
GitHub stars
127
Token cost
~2.7k tokens
SKILL.md length
657 words
Files
12 (incl. scripts)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs.

  • Works in 5 steps: Extract all data → Interpret theme.json → Generate CSS variables → …
  • The user wants to capture a web page screenshot (full page
  • SKILL.md covers When to use this, Quick start, Parameters and Output structure, plus 8 more sections
  • Runs JavaScript scripts from its folder; calls node, npx and npm; reaches npmmirror.com

What it does

Extract Page Data is an agent skill from lintendo/Axhub-Skills. Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs. Use this skill whenever the user wants to capture a web page screenshot (full page or element), extract CSS design tokens or theme information, convert page content to Markdown, collect links and interactive elements from a page, or export a complete page analysis package. Triggers on any URL that isn't an Axure prototype — for Axure URLs use…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `scripts/inject/README.md`, `scripts/inject/extract-links.js` and `scripts/inject/extract-markdown.js`).

It sits in Frontend & Design, covering Design tokens, Schema markup and Browser testing. It works with Playwright. The licence is MIT.

When your agent uses it

  • The user wants to capture a web page screenshot (full page
  • Extract CSS design tokens
  • Theme information
  • Convert page content to Markdown

Example prompts

  • “/extract-page-data”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Extract all data
  2. Interpret theme.json
  3. Generate CSS variables
  4. Restore animations
  5. Handle layered assets

What it can do on your machine

Read from SKILL.md and the folder at commit 1e1c5bc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • npx
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • npmmirror.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extract Page Data loads about 2.7k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:111
    **Connect to an existing Chrome session** (recommended — reuses cookies):

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from lintendo/Axhub-Skills at commit 1e1c5bc, republished under its MIT licence (© lintendo). 657 words, ~2,701 tokens.

Download SKILL.mdSave it as .claude/skills/extract-page-data/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
extract-page-data
description
Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs. Use this skill whenever the user wants to capture a web page screenshot (full page or element), extract CSS design tokens or theme information, convert page content to Markdown, collect links and interactive elements from a page, or export a complete page analysis package. Triggers on any URL that isn't an Axure prototype — for Axure URLs use extract-axure-data instead.

Extract Page Data

Extract structured data from any web page for design analysis, page reconstruction, or content review. The script handles dependency installation automatically — just point it at a URL.

When to use this

  • User shares a web page URL and wants to capture screenshots
  • User asks to extract design tokens, CSS theme, or color palette from a page
  • User wants page content as Markdown text
  • User wants to know what links, buttons, and interactive elements are on a page
  • User asks to export or analyze any web page's structure
  • User wants to take a screenshot of a specific element on a page (use --selector)

Do NOT use for Axure prototypes — use extract-axure-data instead.

Quick start

Run the extraction script from scripts/extract.mjs relative to this skill directory.

bash
# Full page screenshot
node scripts/extract.mjs https://example.com --screenshot

# Screenshot of a specific element
node scripts/extract.mjs https://example.com --screenshot --selector "#hero"

# Extract design tokens (colors, typography, spacing, etc.)
node scripts/extract.mjs https://example.com --theme

# Extract page content as Markdown
node scripts/extract.mjs https://example.com --markdown

# Collect all links and interactive elements
node scripts/extract.mjs https://example.com --links

# Extract everything
node scripts/extract.mjs https://example.com --all

# Export full data pack (screenshot + theme + markdown + links as zip)
node scripts/extract.mjs https://example.com --pack

First run auto-installs Playwright + Chromium to ~/.cache/page-extractor/. This takes 1-2 minutes and happens once.

Parameters

FlagWhat it doesDefault
<url>Web page URL (required)—
--screenshotCapture page screenshotoff
--themeExtract design tokensoff
--markdownConvert page to Markdownoff
--linksCollect interactive elementsoff
--packExport full data pack (zip)off
--allRun all extractionsoff
--selector SELCSS selector to scope extractionwhole page
-o DIROutput directory./page-export
--viewport WxHViewport size1280x720
--wait MSExtra wait after page load (ms)0
--scrollScroll page to trigger lazy contentoff
--scroll-step PXPixels per scroll step800
--scroll-delay MSDelay between scroll steps (ms)200
--no-headlessShow browser windowheadless
--connect-cdp URLAttach to running Chrome (for auth)—
--format FMTScreenshot format: png or jpegpng
--verboseDetailed loggingoff

Output structure

page-export/
├── screenshot.png       # Page or element screenshot    (--screenshot)
├── theme.json           # Design tokens                 (--theme)
├── content.md           # Page text as Markdown          (--markdown)
├── links.json           # Links and interactive elements (--links)
└── page-data.zip        # Full data pack                 (--pack)
    ├── screenshot.png
    ├── theme.json
    ├── content.md
    ├── links.json
    └── meta.json        # Page metadata (title, url, viewport)

Selector support

All extraction modes support --selector to scope to a specific element:

bash
# Screenshot only the header
node scripts/extract.mjs https://example.com --screenshot --selector "header"

# Extract design tokens from a card component
node scripts/extract.mjs https://example.com --theme --selector ".card"

# Get Markdown from the main content area
node scripts/extract.mjs https://example.com --markdown --selector "main"

# Collect links from the navigation
node scripts/extract.mjs https://example.com --links --selector "nav"

Authenticated pages

Some pages require login. Two approaches:

Connect to an existing Chrome session (recommended — reuses cookies):

macOS:

bash
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222

Windows:

powershell
start chrome --remote-debugging-port=9222

Then log in via that Chrome window and run:

bash
node scripts/extract.mjs https://example.com --all --connect-cdp http://localhost:9222

Show the browser window (manual login during extraction):

bash
node scripts/extract.mjs https://example.com --all --no-headless

Dependency installation

The script auto-installs on first run. If it fails:

bash
# Manual install
cd ~/.cache/page-extractor
npm install playwright
npx playwright install chromium

# If Chromium download is slow (e.g. behind GFW):
PLAYWRIGHT_DOWNLOAD_HOST=https://npmmirror.com/mirrors/playwright npx playwright install chromium

Output data formats

theme.json

Design system tokens extracted from computed CSS styles:

json
{
  "colors": {
    "background": [{"value": "rgb(255,255,255)", "count": 42, "tags": ["div"]}],
    "text": [{"value": "rgb(51,51,51)", "count": 28, "tags": ["p","span"]}],
    "border": [{"value": "rgb(221,221,221)", "count": 12}]
  },
  "typography": {
    "families": [{"value": "\"Inter\", sans-serif", "count": 56}],
    "textStyles": [{"size": "14px", "lineHeight": "22px", "weight": "400", "count": 20}]
  },
  "spacing": [{"value": "16px", "count": 15}],
  "radius": [{"value": "4px", "count": 8}],
  "lineWidth": [{"value": "1px", "count": 10}],
  "shadow": {
    "box": [{"value": "0 2px 8px rgba(0,0,0,0.1)", "count": 5}],
    "text": []
  },
  "transitions": [{"value": "all 0.2s ease", "count": 38, "tags": ["button","a"]}],
  "animations": [{"value": "fadeIn|0.3s|ease", "count": 5, "tags": ["div"]}],
  "cssVariables": {
    "--color-primary": "#6366f1",
    "--radius": "0.5rem",
    "--font-sans": "Inter, sans-serif"
  },
  "assets": {
    "backgroundImages": [{"value": "url(hero.webp)", "count": 1, "tags": ["section"]}],
    "images": [
      {"src": "https://example.com/logo.svg", "alt": "Logo", "position": "static", "zIndex": "auto", "siblingImgCount": 1},
      {"src": "https://example.com/hero.png", "alt": "", "position": "absolute", "zIndex": "2", "siblingImgCount": 2}
    ],
    "svgCount": 12
  }
}

Interactive elements on the page:

json
{
  "pageUrl": "https://example.com",
  "pageTitle": "Example",
  "links": [
    {"type": "a", "text": "Home", "href": "https://example.com/", "visible": true},
    {"type": "button", "text": "Sign Up", "visible": true},
    {"type": "form", "href": "https://example.com/login", "visible": true}
  ],
  "totalLinks": 3,
  "visibleLinks": 3
}

Customization

The extraction logic is modular. Each feature lives in its own file under scripts/:

FilePurposeCustomization
lib/browser.mjsPlaywright browser managementChange browser launch args, proxy
lib/screenshot.mjsScreenshot captureModify scroll behavior, format
lib/theme.mjsDesign token extractionAdjust token categories, limits
lib/markdown.mjsHTML → MarkdownChange conversion rules
lib/links.mjsElement collectionAdd new element types
lib/pack.mjsZip packagingChange pack contents
inject/extract-theme.jsBrowser-injected theme logicEdit CSS properties to extract
inject/extract-markdown.jsBrowser-injected markdown logicCustomize text extraction
inject/extract-links.jsBrowser-injected links logicAdd custom selectors
Show full SKILL.md (232 more words)Show less

Theme Clone Workflow

When the user wants to clone or replicate a website's design language, run in this order:

Step 1: Extract all data
bash
node extract.mjs <url> --all --scroll --viewport 1440x900
Step 2: Interpret theme.json
FieldMaps toCSS Variable
colors.background[0]Primary surface--color-bg
colors.text[0]Body text--color-text
colors.text[1]Accent / brand--color-primary
typography.families[0]Main fontGoogle Fonts import
spacing (top 5)Spacing scale--spacing-1 …
radius[0]Default radius--radius-base
transitions[0]Hover timingapply to .card, a, button
cssVariablesDesign tokens (if :root vars exist)use names directly
Step 3: Generate CSS variables
css
:root {
  --color-bg:       <colors.background[0].value>;
  --color-text:     <colors.text[0].value>;
  --color-primary:  <colors.text[1].value>;
  --font-sans:      <typography.families[0].value>;
  --radius-base:    <radius[0].value>;
  --transition:     <transitions[0].value>;
}

If cssVariables is non-empty, prefer those names — they reflect the original design system's intent.

Step 4: Restore animations
  • transitions[0].value → apply as transition: <value> on interactive elements
  • animations → reconstruct @keyframes blocks by animation name; apply to entry elements
Step 5: Handle layered assets

Images with position: absolute and zIndex > 0 in assets.images are overlay layers. Reconstruct with CSS position: relative on the parent and position: absolute on overlays.

Playwright 查缺补漏

导出数据是 best-effort 快照,有时需要回到原始页面补充信息(hover 态样式、懒加载内容、动画效果等)。可使用 Playwright CLI 进行交互式补充采集。

bash
# 打开页面并获取快照
playwright-cli open https://example.com
playwright-cli snapshot

# 交互后采集(如 hover 态)
playwright-cli hover e15
playwright-cli screenshot --filename=hover-state.png

# 获取元素属性
playwright-cli eval "el => getComputedStyle(el).backdropFilter" e5

playwright-cli close

完整命令参考: Playwright CLI 官方技能文档

何时需要 Playwright
场景Playwright 做什么
截图模糊/不完整对指定元素高分辨率截图
样式数据缺失获取完整 computedStyle
hover/focus/active 态触发交互后采集样式
懒加载内容滚动触发后截图
动画/过渡效果定时截图或录屏
响应式断点校验设置 viewport 后截图

Troubleshooting

ProblemFix
Blank screenshotPage needs time to render — add --wait 2000
Missing lazy contentUse --scroll to trigger lazy loading
Need loginUse --connect-cdp (see Authenticated pages above)
Playwright install failsManual install, or set PLAYWRIGHT_DOWNLOAD_HOST
Wrong element selectedCheck selector with browser DevTools first

© lintendo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts) in skills/extract-page-data of lintendo/Axhub-Skills.

  • SKILL.md
  • scripts/extract.mjs
  • scripts/inject/README.md
  • scripts/inject/extract-links.js
  • scripts/inject/extract-markdown.js
  • scripts/inject/extract-theme.js
  • scripts/lib/browser.mjs
  • scripts/lib/links.mjs
  • scripts/lib/markdown.mjs
  • scripts/lib/pack.mjs
  • scripts/lib/screenshot.mjs
  • scripts/lib/theme.mjs

Open the folder on GitHubat commit 1e1c5bc

Compare with similar skills

Extract Page Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extract Page Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extract Page Data this skilllintendo/Axhub-Skills127—~2.7kAutomated safety check: WarnMIT
Extract Design Systemespennilsen/pi122—~2kAutomated safety check: PassMIT
Building With Lobe UIlobehub/lobe-ui2.2k—~2kAutomated safety check: PassMIT
Web Designxiaopu-ai/web-design783—~3kAutomated safety check: PassMIT
Extract Design Systemarvindrk/extract-design-system2361 repos~514Automated safety check: PassMIT
Website Design Taste Extractorsenlindesign/taste-skill381—~5.3kAutomated safety check: PassNone

Similar skills

  • Extract Design System

    espennilsen/pi

    Reverse-engineer a design system from a live website (public URL or localhost).

    122 GitHub stars~2k tokensUpdated 16 days ago
    Frontend & DesignAuto-check passed
  • Building With Lobe UI

    lobehub/lobe-ui

    Build UI with the LobeHub design ecosystem — @lobehub/ui (plus its base-ui, chat, mobile, awesome, brand, mdx, i18n namespaces), @lobehub/icons, @lobehub/charts, @lobehub/fluent-emoji and…

    2.2k GitHub stars~2k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Web Design

    xiaopu-ai/web-design

    Web 视觉设计 SKILL。输入 PRD / 参考 URL / 截图 / 关键词(任意组合),先产出一份标准化 DESIGN.md 设计规范,用户确认后据此生成 UI/UX、视觉、动效、响应式全部达标的 web 代码。专攻 web 端:Landing Page、Portfolio、产品页、博客、个人站、SaaS 介绍页等。当用户说"帮我做个网站""设计一个页面""参考 XX…

    783 GitHub stars~3k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Extract Design System

    arvindrk/extract-design-system

    Extract design primitives from a public website and generate starter token files for your project.

    236 GitHub starsUsed in 1 repo~514 tokens
    Frontend & DesignAuto-check passed
  • Website Design Taste Extractor

    senlindesign/taste-skill

    Reverse-engineers a website's design from its URL into a Markdown and JSON pair holding concrete tokens and the trade-offs behind the look.

    381 GitHub stars~5.3k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Eval Graphics

    elvisun/newsjack

    Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright.

    1.5k GitHub stars~2k tokensUpdated yesterday
    Frontend & DesignAuto-check passed

More from lintendo/Axhub-Skills

All 13 skills in this repo
  • React To Figma Make

    lintendo/Axhub-Skills

    A skill your agent uses when 需要把现有 React、Vite、Next.js、V0 或 AI Studio 页面转换为 Figma Make 可导入的 .fig 资产,或补齐、更新、验证已有 Figma Make 导出壳和 canvas.fig。

    127 GitHub stars~4.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Search Design System

    lintendo/Axhub-Skills

    A skill your agent uses when selecting an existing design system or theme for a product, page, prototype, or implementation through Axhub design knowledge indexes; not when creating or updating a…

    127 GitHub stars~643 tokensUpdated 1 mo ago
    Auto-check passed
  • Axhub Prototype Context

    lintendo/Axhub-Skills

    A skill your agent uses when an annotated prototype URL exposes PRD, directory, annotation, or source context that an AI Agent needs for review, implementation planning, production work, source…

    127 GitHub stars~684 tokensUpdated 1 mo ago
    Auto-check passed
  • Extract Axure Data

    lintendo/Axhub-Skills

    Extract structured data from Axure prototypes using Playwright — screenshots, design tokens, interaction maps, annotations, and page text.

    127 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check: warnings
  • Clone Page

    lintendo/Axhub-Skills

    高精度克隆网页。使用 Playwright 渐进式采集页面数据(DOM 结构、样式、截图、资源), 然后分阶段构建完整的 HTML/CSS 还原。支持快速还原和高精度还原两种模式。

    127 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Figma Content Operator

    lintendo/Axhub-Skills

    用于处理涉及检查、读取、创建、编辑、导出或代码映射的 Figma 内容任务,也用于将 HTML/React/Vue 等网页复制到 Figma,尤其适用于当前选区、设计系统、Figma MCP、Drafito、Figwright 或需要避免 MCP 工具污染上下文的场景。当用户需要可批注 Figma 链接或要求同时绑定工作目录时,也应与 axhub-commentary 组合。

    127 GitHub stars~630 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Extract Page Data

What does Extract Page Data do?

Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs. Extract Page Data is an agent skill from lintendo/Axhub-Skills. Extract structured data from any web page using Playwright — screenshots, design tokens, page text as Markdown, interactive element maps, and full page data packs.

When should I use Extract Page Data?

Extract Page Data fits situations like: the user wants to capture a web page screenshot (full page; extract CSS design tokens; theme information; convert page content to Markdown.

How do I install Extract Page Data in Claude Code?

Run `npx skills add lintendo/Axhub-Skills --skill extract-page-data -a claude-code`. Or copy the skill folder (skills/extract-page-data in lintendo/Axhub-Skills) into .claude/skills/extract-page-data in your project. Claude Code loads it when a task matches its description.

How do I install Extract Page Data in Codex?

Run `npx skills add lintendo/Axhub-Skills --skill extract-page-data -a codex`. Or copy the skill folder (skills/extract-page-data in lintendo/Axhub-Skills) into .agents/skills/extract-page-data in your project. Codex loads it when a task matches its description.

Can I use Extract Page Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lintendo/Axhub-Skills --skill extract-page-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract-page-data, .gemini/skills/extract-page-data, .github/skills/extract-page-data and .opencode/skills/extract-page-data in your project.

What does Extract Page Data need to run?

Going by SKILL.md and its folder, Extract Page Data needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node, npx and npm). Our summary lists: Node.js.

Does Extract Page Data access the network?

SKILL.md names 2 domains. In commands or code: npmmirror.com; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Extract Page Data safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Extract Page Data use?

Extract Page Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extract Page Data use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extract Page Data?

Skills that share tags, products or a category with Extract Page Data: Extract Design System (espennilsen/pi, 122 stars), Building With Lobe UI (lobehub/lobe-ui, 2.2k stars), Web Design (xiaopu-ai/web-design, 783 stars) and Extract Design System (arvindrk/extract-design-system, 236 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extract Page Data?

lintendo (a GitHub user) maintains it in lintendo/Axhub-Skills, which has 127 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 21, 2026.

Source: lintendo/Axhub-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.