Design System
Ohh-889/skyroc
Token architecture, component specifications, and slide generation.
Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design…
$ npx skills add adobe/skills --skill extract -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install adobe/skills extract --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/stardust/skills/extract .claude/skills/extract && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .claude/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extractType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add adobe/skills --skill extract -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install adobe/skills extract --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/stardust/skills/extract .agents/skills/extract && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .agents/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adobe/skills --skill extract -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install adobe/skills extract --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/stardust/skills/extract .cursor/skills/extract && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .cursor/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/adobe/skills.git --path plugins/stardust/skills/extract--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add adobe/skills --skill extract -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install adobe/skills extract --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/stardust/skills/extract .gemini/skills/extract && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .gemini/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install adobe/skills extractInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add adobe/skills --skill extract -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/stardust/skills/extract .github/skills/extract && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .github/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adobe/skills --skill extract -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install adobe/skills extract --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adobe/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/stardust/skills/extract .opencode/skills/extract && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "extract" agent skill from https://github.com/adobe/skills/tree/main/plugins/stardust/skills/extract into .opencode/skills/extract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "extract", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
extractCrawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design…
Extract is an agent skill from adobe/skills. Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an existing site's design, extract or reverse-engineer its design system or brand, capture design tokens from a live site, import a website as the starting point for a redesign, capture the current state before a migration…
Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts (for example `reference/brand-review-template.md`, `reference/brand-surface.md` and `reference/cross-site-sources.md`). Compatibility notes: Requires Node 22+, Playwright with Chromium resolvable from the project, playwright-cli on PATH, and the impeccable skill (github.com/pbakaus/impeccable)…
It sits in Frontend & Design, covering Design tokens, Web scraping and Prototyping. The repository describes itself as: Adobe Skills for Agents. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 985c436. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
nodenpmnpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm and npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Node 22+, Playwright with Chromium resolvable from the project, playwright-cli on PATH, and the impeccable skill (github.com/pbakaus/impeccable) installed alongside stardust.
From compatibility in the SKILL.md frontmatter.
Extract loads about 11k tokens when it runs. Until then it costs about 236 tokens; SKILL.md has 5,123 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from adobe/skills at commit 985c436, republished under its Apache-2.0 licence (© adobe). 5,123 words, ~11,315 tokens.
.claude/skills/extract/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Crawl an existing website, parse each page, extract the brand surface,
and produce a stardust-formatted snapshot of the current state under
stardust/current/. The output describes what the site is; later
sub-commands consume it to decide what it should be.
This skill is descriptive: it does not invent direction, it does not
critique, and it does not modify the live site. It writes only under
stardust/current/ and updates stardust/state.json.
<url> — required. The origin to crawl. Examples: https://example.com,
https://example.com/shop. A path narrows the same-origin crawl to
that subtree.--cap <N> — optional. Override the default 5-page cap. The cap
is intentionally small — a 5-page sample (home + four IA
pillars/templates) is enough for cross-page brand aggregation,
system-component detection, and the brand-review HTML; lift it
(e.g. --cap 25) when a deeper crawl is genuinely needed.--all — optional. Lift the cap entirely; extract every
discovered page after junk filtering. Equivalent to --cap 0.
Use when the user spontaneously asks for a full crawl.--pages <slug,slug,...> — optional. Restrict the crawl to specific
paths (slugs derived per reference/ia-extraction.md). Bypasses
the cap.--refresh <slug> — optional. Re-extract one page that already exists
in state.json.--single — optional. Equivalent to --cap 1. Useful for testing.--wait <fast|medium|spec|auto> — optional. Wait strategy per page.
Default medium. See reference/playwright-recipe.md § Wait modes.--no-junk-filter — optional. Disable the default junk-page filter
in discovery (see reference/ia-extraction.md § Filtering).--no-consent-dismiss — optional. Skip the pre-flight consent /
cookie banner dismissal (see reference/playwright-recipe.md
§ Pre-flight: consent dismissal). Use when the redesign scope
includes the consent surface or the dismissal's side-effects
(script activation that wouldn't otherwise run) must be
avoided. Default is to dismiss, keeping screenshots, voice
aggregation, and per-section style unpolluted by the banner.--dynamics — optional, migration-bound. Record per-page reach
signals of the dynamic surface (data endpoints, forms, modal
triggers, player ids) in each page JSON dynamic section and roll
them up in _crawl-log.json#dynamicSurface. Set by
prepare-migration, replica and migrate's safety net; never by a
bare extract, uplift or audit — dynamics is a migration concern.
Depth and classification belong to the stardust dynamics skill.--concurrency <n> — optional. Parallel browser contexts for the
per-page capture loop. Default 4; sane range 4–8. See
§ Concurrency.--brand-source <url> — optional, repeatable. An additional
same-brand origin whose brand surface enriches the primary
extraction (shallow capture: home + up to 2 nav-linked pages).
See § Cross-site brand sources.--design-source <url> — optional. Design-donor origin: its
design system is captured to stardust/canon-source/ and becomes
the fixed redesign target while the primary origin supplies
content. See § Cross-site brand sources.--prep — optional. Run in migrate-prep mode: lift the cap,
type each page, detect module candidates, capture typed content
slots, emit the prep summary. See § Prep mode below. Typically
invoked via the prepare-migration orchestrator skill rather
than directly.Run the master skill's setup procedure first
(skills/stardust/SKILL.md § Setup): impeccable dep check, context
loader, state read.
Additional checks for this sub-command:
Playwright availability. The extraction step needs a real
browser. Detect Playwright in this order: a Playwright MCP server,
then a project-importable playwright module. The npx playwright probe is NOT sufficient — it confirms the CLI
(which resolves a global install) but the recipe and
scripts/crawl.mjs do import { chromium } from 'playwright',
and ESM module resolution does not honour a global install or
NODE_PATH — the import throws ERR_MODULE_NOT_FOUND even where
npx playwright --version succeeds. Verify the module is
import-resolvable from the project root (probe:
node -e "import('playwright').then(()=>process.exit(0))"); if it
isn't, run npm i -D playwright pixelmatch pngjs cheerio --legacy-peer-deps
(devDependencies, never --no-save — #125) or use the Playwright MCP
server, before crawling. The
--legacy-peer-deps flag is required on aem-boilerplate targets
(their pinned eslint@8 makes a plain npm i exit ERESOLVE
before playwright is even considered). Don't trust the CLI
probe alone.
--no-save installs are ephemeral. Any later real npm i (e.g. a setup step
adding a devDependency) prunes non-manifest packages, silently
removing playwright mid-pipeline. Every downstream skill that
renders (prototype, migrate, deploy, diff) must re-run the
import-resolvability probe — and re-install on failure — at the
start of its own run, not assume extract's install survived.
Script location matters. ESM resolves import 'playwright'
from the script's directory and the plugin tree ships no
node_modules: copy every script byte-identical into the project
(stardust/scripts/…) and run the copy.
Bundled crawler. skills/extract/scripts/crawl.mjs is a
runnable reference implementation of this whole sub-command —
browser config + bot-management fallback, consent dismissal, wait
assets/screenshots/<slug>.png, consumed by the Phase 2.5
vision gate), response validation, and the §
Capture-hygiene hardening (visibility filter, interstitial drop,
SPA-shell flag, modal textContent capture, tracking-pixel
discounting, cross-page duplicate detection). Prefer invoking it over
hand-rolling a Playwright script per run; extend its in-page
capture() to cover any recipe field it doesn't yet emit.Origin collision. If stardust/state.json already records
site.originUrl and the new <url> is a different origin, stop and
ask before clobbering. Stardust does not silently mix two sites in
one project.
Flow guard (migration asks only). If the ask carries migration
intent ("migrate", "to EDS", "re-platform", "1:1", "replica") and
stardust/state.json exists — or is about to be created — without
flow, hand back to the master skill § Two migration flows before
crawling: the flow is chosen and stamped there, and a keep-design
ask enters through replica (which invokes this skill with --prep
itself). A bare extract <url> for a redesign, audit or uplift is
unaffected. (Recorded: extract on a raw URL as the entry of a
same-design migration; the agent then built its own importer beside
replica.)
Browser contexts. Open a fresh BrowserContext per capture
worker (§ Concurrency; default 4). Run the consent dismissal
pre-flight per reference/playwright-recipe.md § Pre-flight:
consent dismissal unless --no-consent-dismiss is set.
Cookies persist within a context but not across contexts —
re-establish consent state per context (re-run the dismissal on
the worker's first page, or clone the probe context's
storageState). Record the resolved method in
_crawl-log.json#consent.method.
Bot-management probe. When the first navigation in the run
returns ERR_HTTP2_PROTOCOL_ERROR or ERR_QUIC_PROTOCOL_ERROR,
or hangs through the entire hard-cap on what should be a fast
origin, do not retry headless — switch to headed Chrome per
§ Failure modes → Bot-management block and record the switch in
_crawl-log.json#discovery.fetchTechnique so re-runs start in
headed mode without rediscovering the issue.
Discover the page inventory before crawling. Procedure in
reference/ia-extraction.md. In summary:
Fetch <origin>/sitemap.xml, then <origin>/sitemap_index.xml,
then check robots.txt for Sitemap: directives.
If no sitemap is reachable, run a same-origin BFS crawl from
<url>, depth-limited to 3, link-extracting from rendered HTML.
Filter the discovered URL list: same origin only, exclude
mailto:, tel:, anchor-only links, query-only variations,
common asset paths (.css, .js, .pdf, image extensions).
De-duplicate trailing-slash variations.
Apply the junk-page filter (reference/ia-extraction.md §
Junk-page filter) unless --no-junk-filter is set. Surface the
filtered list to the user as overridable.
Apply the cap (default 5, or --cap, or --all for no cap)
and proceed silently. Print an informational summary of
what was kept and what was cut — but do not gate on user
confirmation. Users who want different scope set it
spontaneously at command time:
$stardust extract https://example.com # default 5 pages
$stardust extract https://example.com --cap 25 # bump to 25
$stardust extract https://example.com --all # lift the cap
$stardust extract https://example.com --pages home,about,pricing
$stardust extract https://example.com --single # just the entry URLThe agent reads spontaneous scope intent from the user's prompt (e.g. "extract all pages", "look at just the home and pricing", "do a full crawl") and applies the equivalent flag. No re-confirmation needed once intent is clear.
Informational output (not a prompt — proceed immediately):
Discovered 38 pages on https://example.com (sitemap.xml).
Filtered as likely junk (5): /test/, /sample-page/, /holiday1/, ...
Selecting 5 highest-priority pages:
- / (home)
- /about
- /pricing
- /products
- /contact
Cut (28 pages, --all to lift): /blog/post-1, /blog/post-2, ...
Extracting...Selection heuristic: page-type checklist first, then score-based
ranking (home + IA-pillar keywords + sitemap priority − archive /
version markers). See reference/ia-extraction.md § Page
selection and § Priority for the cap. The English-only keyword
list is a known limitation for localized sites.
Write the discovered list to stardust/current/_crawl-log.json
(created if absent) with _provenance and the full discovery
reasoning, including filteredAsJunk[] and userChoice. This is
an audit trail, not a state file.
For each page in the cap-respecting list, render with Playwright
following reference/playwright-recipe.md. Captures run
concurrently per § Concurrency. The recipe is
mandatory per page — in particular, do not skip the wait, scroll, or
capture-list steps:
medium; see § Wait
modes in reference/playwright-recipe.md)prefers-reduced-motion: reducewaitMs and waitMode in the per-page _provenanceCapture per page (full schema in reference/current-state-schema.md):
Page metadata (title, meta description, OG tags, theme-color)
Semantic structure: heading outline, landmark roles, sections
Hero headline + lede (resolved) — heroHeadline / heroLede
picked by font-size × hero-region with a junk/hidden-state filter and
a clean meta-description fallback (per reference/playwright-recipe.md
§ Capture list 5-bis). Required for JS-rendered sites whose
document-order headings surface modal / promo / count junk
instead of the real tagline.
Content: visible text per section (full innerText, no
truncation per reference/playwright-recipe.md § Capture
list 7), structured paragraphs (body[]), lists, FAQ Q/A
pairs, and review/testimonial quotes per
§ Capture list 7-bis. Without these structured fields,
every body region under a heading falls back to placeholder
signature at migrate time.
CTA labels and href targets, link inventory (internal vs external)
Per-section computed style summary: dominant colors, font families in use, spacing rhythm, border-radius, shadows
Media inventory: img with currentSrc/srcset captured with
query strings intact plus a resolves flag (HEAD/GET with browser
UA + Referer), intrinsic dimensions, inline SVG count, video/iframe
presence, cssBackgrounds[] (including pseudo-element ::before/
::after walks per § Capture list 11) so background-image
heroes and motifs do not silently disappear and 404ing CDN
images are flagged before migrate ships about:error.
Font files captured via network-intercept (per § Capture list
16): every woff2/woff/ttf/otf response saved under
assets/fonts/ and recorded in _brand-extraction.json#type.files[]
with licensing flag.
Icon-font detection (per § Capture list 17): when the page
uses [class^="icon-"] with non-default ::before
font-family + codepoint, capture the family, save the file,
and record the iconClass → codepoint table in
_brand-extraction.json#iconFont.
Interactive elements: forms (with field types), buttons, modals detected by ARIA roles
Full-page screenshot to assets/screenshots/<slug>.png —
script-captured by the bundled crawler after the wait/scroll
settle (viewport-only fallback on extremely tall pages; mode in
_signals.screenshotMode, relative path in the page JSON
screenshot field)
Dynamic surface (only with --dynamics) — per-page reach
signals: endpoints, third-party script hosts, forms, modal triggers,
player ids, hydration hints, in the page JSON dynamic section and
_crawl-log.json#dynamicSurface (schema in
reference/current-state-schema.md § Dynamic). Evidence only; the
stardust dynamics sub-skill probes archetypes in depth and decides.
Save to stardust/current/pages/<slug>.json with _provenance as the
first key. The bundled crawler also saves the settled rendered DOM
verbatim as stardust/current/pages/<slug>.html (page.content()
after the wait/scroll settle; path in the record's renderedHtml
field). Capture once, parse offline: every downstream importer or
sibling generator iterates its extraction against this artifact —
free, reproducible, and provenance — instead of re-running live
probes per selector guess (recorded: 4+ live round-trips per page
family before the switch). Live probes stay for what the static DOM
cannot answer: geometry and computed styles. Save referenced media to stardust/current/assets/media/
preserving basename plus a short content hash.
Live-render evidence (synthesis is forbidden). Refuse to mark
a page extracted in state.json unless its _provenance
contains renderedBy: "playwright", an ISO-8601 fetchedAt, a
positive integer waitMs, a waitMode from the recipe, and a
final httpStatus in the 2xx/3xx range. These five fields are
the contract enforced by reference/current-state-schema.md
§ Live-render evidence and read back by every downstream phase
via validateProvenance() per
skills/stardust/reference/state-machine.md § Provenance
validation. Synthesizing a page record from
_brand-extraction.json plus URL patterns plus captured photos
— the 2026-04-30 e-commerce shortcut — is the failure mode this
guard exists to prevent. When the agent (or a delegated sub-
agent) cannot satisfy the contract for a page, treat the page
as a Phase 2 failure: record under _crawl-log.json#crawl.failures[]
with errorClass: "ProvenanceMissing" and continue.
Mark the page extracted in state.json immediately after each
successful page write. If a page fails, record the error in
_crawl-log.json and continue — extraction is best-effort per page.
Before anything downstream is authored, look at each captured
page's screenshot (assets/screenshots/<slug>.png) — the multimodal
model reads the image — and verify it against the extracted record:
cssBackgrounds: [] record look believable, or is imagery
visibly present — a silent capture failure?Read the whole page through its thumbnail —
node stardust/scripts/thumb.mjs stardust/current/assets/screenshots/<slug>.png --width 480
(skills/extract/scripts/thumb.mjs, copied beside crawl.mjs; a
box-filter downscale, so 1-px rules and hairline borders survive; its
--max-bytes cap, 150 KB by default, re-encodes a heavy page narrower
first and crops only as a last resort, never below 60 % of the page) —
never the full-resolution screenshot, and read any detail as a crop of
the full capture. The check is of the WHOLE page: read the script's
stdout line before the image, and when it prints a share below 100
(cropped at <n>px of <H> = <share>%) run the same command with
--offset <n> and read <slug>-thumb-<n>.png too, repeating while its
rows <n>-<end>px line ends short of <H>. A recorded hands-off run's
thumbnails kept a median 28 % of each page and the check read them as
the page. A thumbnail enters the context only after thumb.mjs has
written it (≤ 150 KB, or the floor thumbnail it named on stderr, exit 1)
— or the vision check runs in a subagent that reads the thumbnails and
returns one line per page
(<slug>: ok | recaptured | suspect — <note>); a recorded hands-off run
read eight *-thumb.png files of 268–608 KB each into one context —
3.2 MB of pixels for eight verdicts.
On mismatch, re-run that page's capture with the escalation ladder
before proceeding: bump the wait mode one step
(reference/playwright-recipe.md § Wait modes), then headed Chrome
(§ Bot-management fallback), then a fresh browser context. Record
the outcome per page in _crawl-log.json#visionCheck[]:
{ "slug": "pricing", "verdict": "recaptured", "notes": "record said zero CSS backgrounds; screenshot shows a full-bleed photo hero" }verdict is "ok" | "recaptured" | "suspect" — suspect means the
mismatch survived the ladder; downstream phases treat that record as
unreliable. Vision is the authoritative capture check; the
heuristic defenses (low-media flag, spaShellSuspect, duplicate
hash) remain as cheap early signals but no longer gate alone.
Run after the capture phases (2–2.5). Aggregation may proceed
incrementally as concurrent page captures complete (§ Concurrency),
but the written file must reflect every extracted page — including
brand-source pages per § Cross-site brand sources.
Produces stardust/current/_brand-extraction.json
per reference/brand-surface.md. Some fields are home-only (logo,
voice samples, register heuristic); the visual tokens that drive
DESIGN.md (palette, radius, shadow, type) are aggregated across all
extracted pages to avoid the home-page bias documented in
brand-surface.md § Aggregation scope.
Computed-style census — run the shipped instrument; author no
probe. node <plugin>/skills/extract/scripts/style-census.mjs
(copied into the project like crawl.mjs, § Setup) measures every
captured page at 1440 and writes
stardust/current/_computed-styles.json; add --width 360 when the
breakpoints include it. One more live pass over every page — run it
ONCE, after the crawl, in the background; it opens the live side
through live-session.mjs like every other live instrument. Read the
palette, type, motif and hover values from the file's aggregate and
cite each from its sources[] (url, selector, prop); the
per-page detail sits under pages[url][width].
Content-cap model — run the shipped probe; author no lift (#124).
node stardust/scripts/replica/cap-probe.mjs <archetype-url>… --write-design stardust/current/DESIGN.json (replica's script; copy it with live-session.mjs)
renders each archetype at 1440 and at a DERIVED wide width —
max(2560, largest cap × 1.25); a probe AT a cap's width cannot see it — and
records every box that stops growing (kind shell / content / module,
px, authored via, tier) into extensions.breakpoints: containerMaxWidth
(the --max-width source; null when fluid), probeWidth, caps[],
modules[]. One navigation per archetype after the census (never beside
it), once DESIGN.json is authored (it merges). A cap on the shell or
on every archetype is design intent whatever its value; capRegister lists
single-module or archetype-divergent caps for the inconsistency register.
The cap model never records layout breakpoints: a redesign target keeps its own two;
replica keeps the source's, or maps them to --target-breakpoints.
Captures:
<img> with
logo-ish class/id → apple-touch-icon → og:image → favicon →
synthesized placeholder. Save to stardust/current/assets/logo.<ext>.stardust/current/assets/favicon.<ext>, per
reference/playwright-recipe.md § Favicon capture. Downstream,
prototype embeds it in the proposed page head and deploy ships
it to the Edge Delivery site.brand-surface.md § Modular-scale
audit) and emit scaleAudit.kind = "modular" | "ad-hoc"._provenance.notes.direct later but extracted now so the network round-trip is over.voice.heroImage (per reference/brand-surface.md
§ heroImage resolution), so downstream prototype picks the
live hero rather than the og:image from the raw media list.<video> / HLS / canvas / WebGL /
Lottie / animated SVG / scroll-driven motion), elevate it to
voice.heroMedium (per reference/brand-surface.md § heroMedium
resolution). This is the page's signature; without elevation
downstream prototype flattens it to a static hero. A non-null
heroMedium triggers signature
preservation (skills/stardust/reference/intent-dimensions.md
§ 8b) at prototype time.reference/playwright-recipe.md
§ Capture list 17, populate _brand-extraction.json#iconFont
with family, file path, and the iconClass → codepoint
table so prototypes can render the brand's actual icons.reference/brand-surface.md § System components. Required —
these are usually the most load-bearing surfaces and must not
silently disappear from the redesign target._brand-extraction.json#origins[] per reference/brand-surface.md
§ Origins. Single-origin runs emit a one-entry array; brand-source
runs attribute widened evidence per origin.Do not invent values. Every captured value cites a source selector or
URL in _brand-extraction.json for traceability.
stardust/current/PRODUCT.md and DESIGN.mdThe current-state PRODUCT.md and DESIGN.md are descriptive, not
authored — there is no interview to run because the user is not
defining intent here, the agent is describing the existing site. Write
them directly using impeccable's format specs (impeccable's skill
directory is what node <plugin>/skills/stardust/scripts/impeccable-version-check.mjs --where
prints; read each spec by section, not whole):
reference/init.md § Write PRODUCT.md. File order: stardust's
<!-- stardust:provenance … --> block first (per
../stardust/reference/artifact-map.md § Provenance, as for every
file this skill writes), then # Product, then the
<!-- impeccable:product-schema 1 --> comment verbatim, then
Platform (web), Users, Product Purpose, Positioning,
Capabilities and Constraints, Brand Commitments, Evidence on Hand, Product Principles, Accessibility & Inclusion; omit a
section rather than pad it. Under Brand Commitments record the
register guess from the brand surface (sites that read as
marketing/landing → brand; tools/dashboards → product; ambiguous
→ brand with a note), the observed brand personality and the
observed anti-references. Evidence on Hand lists what was captured
under stardust/current/ with paths. Populate Users, Product Purpose, Positioning and Product Principles from the captured
copy and the brand surface. Where the agent must infer, mark the
section with _provenance: inferred and a one-line basis sentence.reference/document.md. Populate frontmatter
(colors, typography, rounded, spacing, components) from
the captured tokens. The extensions block of DESIGN.json carries
v1's componentStyle, motifs, and voice arrays so nothing is
lost, and breakpoints — the cap-probe sizing model — verbatim.Stardust does not invoke $impeccable init (formerly teach) or
$impeccable document for the current-state files: those commands write to project
root (the target) and run an interview. Stardust authors the
descriptive snapshot directly. The format spec from impeccable is the
contract; the runtime command is not.
The target-state PRODUCT.md and DESIGN.md at the project root are
written by $stardust direct in Phase 2 of the pipeline, not here.
stardust/current/brand-review.htmlAfter Phase 4 writes the descriptive PRODUCT.md and DESIGN.md, emit
the current-state brand review per
reference/brand-review-template.md.
The brand-review HTML is the first surface a human can eyeball to verify the extraction before committing to a redesign direction — misreads that are invisible in JSON (a wrong dominant radius, a missing system component, a single-page palette bias) are obvious to the eye while they are still cheap to fix (re-extract is fast; re-direct + re-prototype is not).
The template is mandatory. In particular:
reference/brand-review-template.md § Detectors. Each rule is
mechanical; emit a tension card whenever the trigger condition
matches. The review may ship with zero tensions if the data is
too thin to evaluate, but the detectors must always be run._brand-extraction.json § type under Typography).If the data for a section is missing, omit the section — do not fabricate placeholders. The coverage callout at the top reflects what is missing.
After all Phase 2-5 writes succeed:
Update stardust/state.json (schema in
skills/stardust/reference/state-machine.md):
site.originUrl, site.extractedAt, site.pageCap,
site.totalDiscovered, site.crawledpages[] — one entry per crawled page with status: "extracted",
filled currentStatePath, empty prototypePath and migratedPathdesignSource — only when --design-source was used:
{ url, capturedAt, path } per § Cross-site brand sourcesPrint a one-screen summary:
Extracted https://example.com (5/38 pages, sitemap.xml)
stardust/current/
PRODUCT.md, DESIGN.md, DESIGN.json, brand-review.html,
pages/ (5), assets/logo.svg, _brand-extraction.json, _crawl-log.json
Per-page evidence:
slug live waitMode waitMs status media(img/bg)
/ yes medium 2380 200 38/6
/pricing yes medium 1940 200 9/0 ⚠ low-media
/contact yes domcontentloaded(fb) 8000 200 3/0
...
Wait summary: 4 resolved at medium (avg 2.4s), 1 fallback (timed out at 8s)
→ /contact may be under-captured; consider --refresh
Media summary: 1 page flagged low-media (/pricing) — see media-coverage check
Vision check: 4 ok, 1 recaptured (/pricing) — _crawl-log.json#visionCheck
Open stardust/current/brand-review.html to verify the extraction
before running $stardust direct.
Coverage note: extracted 5 of 38 discovered pages; 5 pages
covering distinct templates is usually sufficient for the
cross-page brand surface. Widen with --cap <N> or --pages.
Next: $stardust direct (resolve a redesign direction)The per-page evidence table is mandatory. The live column
is yes when _provenance.renderedBy === "playwright" AND
waitMs > 0, else no. A no row means the page record was
not produced by a live Playwright render — the visible column
is the defense-in-depth signal for the failure mode the
write-time guard exists to prevent (2026-04-30 e-commerce run). A
maintainer scanning the summary should see yes on every row.
Compute the wait summary by grouping each page's _provenance.waitMode
and averaging waitMs. List slugs whose waitMode ends in
(fallback) (rendered as (fb) in the table for width) as
candidates for --refresh.
The media(img/bg) column is the analogous defense-in-depth
signal for imagery. It prints <count of media.imgs> / <count of media.cssBackgrounds>. Flag a row ⚠ low-media when the page
reads as brand/marketing (register brand, or a landing/solution/
product template) yet has cssBackgrounds: [] and few large
rasters (no media.imgs entry with intrinsic width ≥ 600). That
combination is the signature of a silently-failed background /
lazy-media walk — a capture pass that specs the full background
walk (playwright-recipe.md § Capture list 11) yet silently
produces nothing still ships an image-less capture (2026-06-26
a SaaS site: cssBackgrounds: [] on every page, all product
imagery lost). A flagged row is the cue to re-run that page with
--refresh (and, if it persists, to fall back to headed Chrome per
§ Bot-management fallback). A maintainer scanning the summary should
treat a brand-register site with all-zero bg counts as suspect,
not as "this site uses no background images."
Two flags widen extraction beyond the primary origin — read
reference/cross-site-sources.md in full whenever either flag is
present. Both are opt-in; without them this section is inert.
Core contract (merge rules and capture shapes in the reference):
--brand-source <url> (repeatable) — a same-brand sibling
gets a shallow capture (home + ≤2 nav-linked pages, full recipe +
provenance) under stardust/current/brand-sources/<host>/.
Evidence, not inventory: never enters state.json.pages[].
Its palette/type/motif/voice evidence aggregates into
_brand-extraction.json with per-origin attribution
(origins[]); conflicts resolve toward the primary; widened
traits are attributed, never invented. Excluded from system
components, voiceTable, and cross-promo detection.--design-source <url> — a design donor is captured to
stardust/canon-source/ (same shapes as current/, default cap),
its descriptive DESIGN.md/json derived per Phase 4 rules, and
state.json.designSource = { url, capturedAt, path: "stardust/canon-source/" } stamped. direct pins the donor
system as the target (its § Mode A); donor evidence never
aggregates into the primary _brand-extraction.json.After the primary crawl (Phases 2–2.5), harvest candidate same-brand origins from evidence already captured — no extra navigation:
hreflang alternates on other domainsshop., careers., country subdomains)og:site_name matches across captured pagesList candidates with confidence + the evidence line in the crawl
report and in _crawl-log.json#siblingCandidates[]:
{ "origin": "https://example.co.uk", "confidence": "high", "evidence": "hreflang alternate on 4/5 pages", "decision": "included" }--brand-source?" — and proceed on the answer.state.json.handsOff is true): auto-include up to
2 high-confidence candidates as brand-sources; record the
decision in siblingCandidates[].decision.Discovery is capped and cheap: harvesting reads captured evidence only, and each included sibling gets the shallow brand-source capture (≤ 3 pages). It must never balloon the crawl.
| Path | Purpose |
|---|---|
stardust/current/PRODUCT.md | Descriptive strategy of the existing site (impeccable format) |
stardust/current/DESIGN.md | Descriptive visual system (Stitch format) |
stardust/current/DESIGN.json | Sidecar with extensions for motifs, voice, components |
stardust/current/brand-review.html | Self-contained visual review of the extraction (first eyeball-able artifact) |
stardust/current/pages/<slug>.json | Per-page parsed structure + content |
stardust/current/pages/<slug>.html | Settled rendered DOM (crawler sidecar; parse offline, never re-scrape) |
stardust/current/assets/logo.<ext> | Extracted logo |
stardust/current/assets/favicon.<ext> | Site favicon (first-class asset; prototype head + deploy consume it) |
stardust/current/assets/media/ | Extracted media referenced by pages |
stardust/current/assets/screenshots/ | Per-page full-page screenshots, script-captured by crawl.mjs (Phase 2.5 vision gate + brand-review) |
stardust/current/_brand-extraction.json | Consolidated brand surface (palette, type, motifs, voice, system components) |
stardust/current/_crawl-log.json | Discovery + crawl audit trail (incl. visionCheck[], siblingCandidates[]; dynamicSurface reach roll-up only with --dynamics) |
stardust/current/brand-sources/<host>/ | Shallow same-brand captures (only with --brand-source) |
stardust/canon-source/ | Design-donor capture + descriptive DESIGN.md/json (only with --design-source) |
stardust/state.json | Updated with site + per-page status (+ designSource stamp) |
Page captures run concurrently: the Phase 2 queue is drained by
4–8 parallel browser contexts (--concurrency, default 4). Each
worker owns its BrowserContext; consent state is re-established per
context (Setup step 3). Media resolves / HEAD checks are batched
with Promise.all. Brand-surface aggregation (Phase 3) may proceed
incrementally as pages complete, so long as the written
_brand-extraction.json reflects every extracted page. The bundled
crawler implements the pool (crawl.mjs --concurrency <n>).
Across processes, per state-machine.md: stardust does not lock. Two
concurrent extracts on the same project are last-write-wins. Document
this in the user report; do not engineer around it.
_crawl-log.json,
end with a partial state. State.json reflects only successfully
extracted pages. User can re-run; already-extracted pages are
skipped unless --refresh <slug>.reference/playwright-recipe.md § Response validation. Each
produces a distinct error class (HTTPError, ContentTypeError,
EmptyPageError) recorded in _crawl-log.json#crawl.failures[].
Failed pages do not appear in state.json as extracted —
they appear only in the failure log. Without this validation a 5xx
page silently lands as an empty success and propagates wrong data
to direct and prototype.ERR_HTTP2_PROTOCOL_ERROR,
ERR_QUIC_PROTOCOL_ERROR, or hangs through the hard-cap on a
TLS/H2 fingerprint check, the issue is JA3/H2 fingerprinting on
bundled-chromium-default headless mode — not auth, not network.
Switch to headless: false, channel: 'chrome' per
reference/playwright-recipe.md § Bot-management fallback. Do
not retry headless: it will fail identically. The headed
fallback works against most enterprise / commerce origins;
playwright-extra + stealth plugin is a non-standard escape
hatch for the residual cases. The headed window pops visibly,
which is acceptable for interactive runs and unacceptable for
unattended pipelines — surface this to the user when first
triggered. Note for asset harvest: a page-level bot wall usually
does NOT gate assets — media/CSS/font URLs commonly return 200 to
a plain browser-UA curl even while every page navigation is
challenged (Cloudflare challenge, field-confirmed). Probe one
asset with curl BEFORE reaching for in-page-fetch machinery; the
in-page harvest is the fallback, not the default.reference/playwright-recipe.md § Wait modes), fall back to
domcontentloaded and capture what is rendered. Record the
fallback in the per-page _provenance.waitMode and surface in the
wait-summary line of the final report.errorClass: "ProvenanceMissing" in
_crawl-log.json#crawl.failures[]) and continue. Synthesizing
a page record from _brand-extraction.json plus URL patterns
plus captured photos at "semantically matching" template
positions is forbidden. The shortcut produces output
indistinguishable from a successful run and propagates
fabricated content through every downstream phase
(2026-04-30 e-commerce run: 20 of 25 pages synthesized, caught
four phases later).When invoked with --prep, extract runs an extended pass that
prepares the inventory for migration — read
reference/prep-mode.md in full before running any --prep
extraction. Discovery-mode runs (without --prep) are unchanged:
small cap, no typing, no module detection. The flag is intended for
the prepare-migration orchestrator, though direct invocation is
supported.
Core contract (procedure, formats, and detection rules in the reference):
--prep implies --all — migration coverage requires the full
junk-filtered inventory, not the discovery cap.landing | article | listing | program | form | static | unique) are LLM-inferred into state.json.pages[].type
(discovery-mode runs leave type null); the user confirms in
direct --prep.DESIGN.json.extensions.modules[] with status: "candidate";
per-page JSON gains a typed slots section per page-type.crawl.mjs takes every captured locale root (a one-
or two-segment path of locale codes — /en, /fr-ca, /ca/en),
lists the same-origin targets linked from that page under its own
path, and marks the ones missing from the capture:
_crawl-log.json#captureGaps (roots[] with root, slug,
linked, uncaptured[], plus detail, a ready ledger string) and
one [crawl] capture gaps: stderr line. Copy roots[] to
state.json.site.captureGaps (the site block is extract's; an
uncaptured target never becomes a pages[] row) and carry detail
verbatim in the ledger end line that closes the extraction —
ledger.mjs extract prep end --detail "… uncaptured first-level targets: ca/en 15, fr/fr 0" (replica's extract end repeats it) — so
the owner decides on the gap now (widen with --pages, or record it
as scope debt), not in the rollout link audit: two recorded hands-off
runs captured /ca/en as a locale root while 15 /ca/en/* pages
existed on the source, and found out six hours later, when 15 links
answered 404 and the gap exceeded the 12-page deliver threshold.Provenance: <live>/<total> live line is mandatory, and any
ratio short of <total>/<total> means the run failed the
synthesis guard and is incomplete.reference/prep-mode.md § Sub-agent prompt requirements);
missing any of the three is a recipe violation.reference/playwright-recipe.md — viewport, capture list, logo locator chain.reference/ia-extraction.md — sitemap + BFS crawl + cap procedure.reference/current-state-schema.md — per-page JSON schema.reference/brand-surface.md — consolidated brand-surface schema.reference/brand-review-template.md — current-state brand-review HTML contract + Tensions detectors.reference/prep-mode.md — full --prep procedure (typing, module candidates, typed slots, prep summary, sub-agent requirements).reference/cross-site-sources.md — full --brand-source / --design-source procedure (shallow capture, merge rules, canon-source donor).skills/stardust/reference/state-machine.md — state.json contract.skills/stardust/reference/artifact-map.md — provenance shape.© adobe, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (scripts) in plugins/stardust/skills/extract of adobe/skills.
Open the folder on GitHubat commit 985c436
Extract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Extract this skilladobe/skills | 196 | — | ~11k | Automated safety check: Pass | Apache-2.0 | |
| Design SystemOhh-889/skyroc | 795 | 11 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Stitch Taste Design Systemgoogle-labs-code/stitch-skills | 8.4k | 15 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Scalar Design Systemscalar/scalar | 16k | — | ~2.7k | Automated safety check: Pass | MIT | |
| Tailwind V4 Shadcnever-works/ever-works | 158 | 2 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Using Ouds Web Version Major MinorOrange-OpenSource/Orange-Boosted-Bootstrap | 222 | — | ~4.2k | Automated safety check: Pass | MIT |
Ohh-889/skyroc
Token architecture, component specifications, and slide generation.
google-labs-code/stitch-skills
Generates a DESIGN.md design-language file for Google Stitch that encodes color, typography, layout, component behavior and motion rules to avoid generic AI-looking UI.
scalar/scalar
Scalar's design system — design tokens, theming (@scalar/themes), CSS variables, and the @scalar/components library.
ever-works/ever-works
Production-tested setup for Tailwind CSS v4 with shadcn/ui, Vite, and React.
Orange-OpenSource/Orange-Boosted-Bootstrap
Provides comprehensive knowledge of the OUDS Web library (Orange Unified Design System for Web), a Bootstrap-based CSS/JS framework for building Orange-branded web interfaces.
MaxHan7/frontend-ui-standards-skill
A skill your agent uses when implementing, refactoring, or reviewing frontend UI across SwiftUI, React, React Native, Flutter, web, or mobile apps.
adobe/skills
Scaffolds, implements, deploys and debugs Adobe Runtime actions in App Builder projects, with templates for webhooks, events, database CRUD, sequences and Asset Compute workers.
adobe/skills
Launches Chrome with an unpacked extension over CDP, opens its sidepanel, popup or options page, and hands over to cdp-connect for clicks, typing and screenshots.
adobe/skills
Extracts icons, metadata, text, forms, videos and social links from any web page with playwright-cli, with SVG icon classification and cleanup.
adobe/skills
Detect all languages used on a webpage — both declared (html@lang, hreflang alternate links, nested lang= attributes, meta content-language) and actually present in the body text (Google CLD3 via…
adobe/skills
Prepare any webpage for clean interaction by detecting and removing disruptive overlays (cookie banners, GDPR consent, modals, popups, newsletter signups, paywalls, login walls).
adobe/skills
Reduce a webpage to a structural skeleton with semantic tokens.
Categories
Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design…. Extract is an agent skill from adobe/skills.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site.
Extract fits situations like: the user wants to analyze an existing sites design; reverse-engineer its design system; capture design tokens from a live site; import a website as the starting point for a redesign.
Run `npx skills add adobe/skills --skill extract -a claude-code`. Or copy the skill folder (plugins/stardust/skills/extract in adobe/skills) into .claude/skills/extract in your project. Claude Code loads it when a task matches its description.
Run `npx skills add adobe/skills --skill extract -a codex`. Or copy the skill folder (plugins/stardust/skills/extract in adobe/skills) into .agents/skills/extract in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adobe/skills --skill extract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extract, .gemini/skills/extract, .github/skills/extract and .opencode/skills/extract in your project.
Going by SKILL.md and its folder, Extract needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node, npm and npx). Our summary lists: Node.js. Compatibility (from SKILL.md): Requires Node 22+, Playwright with Chromium resolvable from the project, playwright-cli on PATH, and the impeccable skill (github.com/pbakaus/impeccable) installed alongside stardust..
SKILL.md contains no URLs. Its commands use npm and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Extract is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 11k tokens (SKILL.md is roughly 45k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Extract: Design System (Ohh-889/skyroc, 795 stars), Stitch Taste Design System (google-labs-code/stitch-skills, 8.4k stars), Scalar Design System (scalar/scalar, 16k stars) and Tailwind V4 Shadcn (ever-works/ever-works, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
adobe (a GitHub organization) maintains it in adobe/skills, which has 196 GitHub stars. The repository holds 65 skills in this directory. The repository was last updated on October 7, 2026.
Source: adobe/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.