Agent skill

Visual To Spec

by Jason904 in Jason904/ui-skill-lab

A skill your agent uses when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation…

MITAuto-check passedFrontend & Design

Install Visual To Spec

skills CLI
$ npx skills add Jason904/ui-skill-lab --skill visual-to-spec -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jason904/ui-skill-lab visual-to-spec --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jason904/ui-skill-lab.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/visual-to-spec .claude/skills/visual-to-spec && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
visual-to-spec
GitHub stars
144
Token cost
~1.6k tokens
SKILL.md length
694 words
Files
12 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation…

  • Works in 12 steps: Every extracted value must include… → reference-visible must include bbox or… → Phase 1 must output uiType, one of:… → …
  • A user provides a reference UI screenshot and asks to extract visual specs
  • SKILL.md covers Trigger Conditions, Required Inputs, Outputs and Source of Truth, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Visual To Spec is an agent skill from Jason904/ui-skill-lab. Use when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation contract, including dashboards, landing pages, mobile screens, image-led UIs, panorama scenes, 3D exhibit UIs, and screenshot-to-spec tasks.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `README.md`, `examples/component-tree.example.md` and `examples/human-review-needed.example.md`).

It sits in Frontend & Design, covering Design tokens, Landing pages and Mobile UI design. The repository describes itself as: AI-driven UI development skill pipeline — test lab for Codex skills covering PRD-to-pixel workflow. The licence is MIT.

When your agent uses it

  • A user provides a reference UI screenshot and asks to extract visual specs
  • Generate design tokens
  • Create a component tree
  • Produce a frontend implementation contract

Example prompts

  • “/visual-to-spec”

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Every extracted value must include source, confidence, confidenceLabel, and strictness.
  2. reference-visible must include bbox or reference bbox within the same component detail; visibility claims without bbox are invalid.
  3. Phase 1 must output uiType, one of: application-dashboard, panorama-scene, image-led-landing, mobile-app, poster-like-ui, unknown.
  4. If uiType is panorama-scene or image-led-landing, phases 5/6 must use the image layer + overlay layer layout model.
  5. Components not present in the screenshot must be listed in Rejected Assumptions, not placed in the component tree.
  6. Values with confidence < 0.6 must appear in human-review-needed.md; low-confidence layout type must go into human review even if…
  7. Every color in the component tree must exist in tokens.json.
  8. Every component dimension must match layout-spec.md.
  9. Sum of sibling component widths must not exceed parent container width; overlay components must verify bbox does not overflow.
  10. Typography scale must decrease monotonically: h1 > h2 > h3 > body.
  11. No orphan tokens (defined but unused).
  12. No missing tokens (used in component tree but undefined).

What it can do on your machine

Read from SKILL.md and the folder at commit a71b4e2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Visual To Spec loads about 1.6k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 694 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jason904/ui-skill-lab at commit a71b4e2, republished under its MIT licence (© Jason904). 694 words, ~1,629 tokens.

Download SKILL.mdSave it as .claude/skills/visual-to-spec/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
visual-to-spec
description
Use when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation contract, including dashboards, landing pages, mobile screens, image-led UIs, panorama scenes, 3D exhibit UIs, and screenshot-to-spec tasks.
argument-hint
[reference-image]

Version: 2.0.0

visual-to-spec

Convert a finalized UI screenshot into a complete frontend implementation contract.

Trigger Conditions

  • User provides a UI screenshot and asks to extract visual specs
  • User asks to generate design tokens from a screenshot
  • User asks to build a component tree or frontend implementation contract from a screenshot
  • User asks to analyze screenshot layout, color, typography, spacing, overlay controls, or panorama/3D scene UI
  • Keywords: "extract spec from screenshot", "image to design spec", "screenshot to spec"

Required Inputs

ItemDescription
reference.pngFinalized UI screenshot (the only required input)
Project token standardsOptional, adapt to existing standards if available

Outputs

All files written to 03_visual_spec/ directory:

FileContent
visual-analysis.mdUI type, visible element evidence, composition, color, typography, spacing analysis
layout-spec.mdLayout positioning and sizing spec, branching by UI type
component-tree.mdHierarchical component tree with per-component style summaries
tokens.jsonDTCG-inspired design tokens with source/confidence annotations
implementation-risks.mdImplementation risk assessment
human-review-needed.mdLow-confidence items and manual review checklist
DESIGN.mdSummary document referencing all above files

See references/output-files.md for detailed output formats.

Source of Truth

reference.png is the sole visual source of truth.

  • Do not redesign.
  • Do not treat UI elements invisible in the screenshot as screenshot facts.
  • Do not introduce colors, fonts, spacing, panels, charts, or navigation structures absent from the image.
  • Do not default to a dashboard / SaaS / three-column layout; UI type determination must come first.
  • reference-visible components must have bbox and visual evidence in the screenshot; components without bbox must not be labeled reference-visible.
  • Inferred implementation wrapper components (e.g., AppShell, MainContent, SceneImageLayer, OverlayLayer) are permitted but must be labeled source: inferred-implementation.
  • For image-led, panorama-scene, 3D exhibit, and product-render screenshots, output an asset strategy: which visual content must be raster/image assets, and which can be implemented as HTML/CSS/SVG.

All values must carry source and confidence annotations. See references/source-strictness-rules.md for complete rules.

Phase Overview (9 phases, sequential)

#PhaseKey Output
1Global composition analysis + UI type determination + visible element evidencevisual-analysis.md §1
2Color extractionvisual-analysis.md §2 + tokens.json (color)
3Typography extractionvisual-analysis.md §3 + tokens.json (typography)
4Spacing and grid systemvisual-analysis.md §4 + tokens.json (spacing)
5Layout specificationlayout-spec.md
6Component tree constructioncomponent-tree.md
7Token merge and validationtokens.json (final)
8Implementation risk assessmentimplementation-risks.md
9Human review checklisthuman-review-needed.md

See references/extraction-workflow.md for detailed workflow.

Show full SKILL.md (321 more words)Show less

Hard Rules

  1. Every extracted value must include source, confidence, confidenceLabel, and strictness.
  2. reference-visible must include bbox or reference bbox within the same component detail; visibility claims without bbox are invalid.
  3. Phase 1 must output uiType, one of: application-dashboard, panorama-scene, image-led-landing, mobile-app, poster-like-ui, unknown.
  4. If uiType is panorama-scene or image-led-landing, phases 5/6 must use the image layer + overlay layer layout model.
  5. Components not present in the screenshot must be listed in Rejected Assumptions, not placed in the component tree.
  6. Values with confidence < 0.6 must appear in human-review-needed.md; low-confidence layout type must go into human review even if confidence >= 0.6.
  7. Every color in the component tree must exist in tokens.json.
  8. Every component dimension must match layout-spec.md.
  9. Sum of sibling component widths must not exceed parent container width; overlay components must verify bbox does not overflow.
  10. Typography scale must decrease monotonically: h1 > h2 > h3 > body.
  11. No orphan tokens (defined but unused).
  12. No missing tokens (used in component tree but undefined).

Capability Boundaries

See references/script-backed-boundary.md for detailed boundary definition.

CategoryExamples
LLM can doVisual recognition, UI type determination, color estimation, layout judgment, component classification, proportion estimation
Script-assistedPrecise color values, WCAG contrast, delta-E, exact pixel measurement, bbox measurement, visual regression comparison
Human confirmationLow-confidence colors, ambiguous component boundaries, uncertain font identification, whether the main visual needs real assets

When scripts are unavailable, LLM does best-effort estimation and annotates confidence and notes.

Supporting Files Index

references/
  • extraction-workflow.md — 9-phase detailed workflow, UI type branching, asset strategy rules
  • token-schema.md — Token JSON schema definition and field descriptions
  • output-files.md — Structure templates for each output file
  • source-strictness-rules.md — Complete rules for source/confidence/strictness/bbox labels
  • script-backed-boundary.md — Boundary definitions for LLM / script / human operations
examples/
  • tokens.example.json — Complete tokens.json example
  • component-tree.example.md — Component tree example, including panorama-scene vs dashboard branching
  • layout-spec.example.md — Layout spec example, including overlay coordinate model
  • visual-analysis.example.md — Visual analysis example, including UI type, visible elements, rejected assumptions
  • human-review-needed.example.md — Human review checklist example

© Jason904, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in skills/visual-to-spec of Jason904/ui-skill-lab.

  • SKILL.md
  • README.md
  • examples/component-tree.example.md
  • examples/human-review-needed.example.md
  • examples/layout-spec.example.md
  • examples/tokens.example.json
  • examples/visual-analysis.example.md
  • references/extraction-workflow.md
  • references/output-files.md
  • references/script-backed-boundary.md
  • references/source-strictness-rules.md
  • references/token-schema.md

Open the folder on GitHubat commit a71b4e2

Compare with similar skills

Visual To Spec next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Visual To Spec compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Visual To Spec this skillJason904/ui-skill-lab144—~1.6kAutomated safety check: PassMIT
Web Designxiaopu-ai/web-design783—~3kAutomated safety check: PassMIT
Refactoring UIs0xDk/refactoring-ui-skill595—~3.5kAutomated safety check: PassMIT
Web Designdrewnekota/cetus146—~2.1kAutomated safety check: PassMIT
iOS Design GuidelinesMacMagazine/app-iOS170—~767Automated safety check: PassNone
Atlaspacifio/ui154—~2.5kAutomated safety check: PassMIT

Similar skills

  • Web Design

    xiaopu-ai/web-design

    Web 视觉设计 SKILL。输入 PRD / 参考 URL / 截图 / 关键词(任意组合),先产出一份标准化 DESIGN.md 设计规范,用户确认后据此生成 UI/UX、视觉、动效、响应式全部达标的 web 代码。专攻 web 端:Landing Page、Portfolio、产品页、博客、个人站、SaaS 介绍页等。当用户说"帮我做个网站""设计一个页面""参考 XX…

    783 GitHub stars~3k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Refactoring UI

    s0xDk/refactoring-ui-skill

    Design and improve user interfaces using the concrete rules from Refactoring UI (Wathan & Schoger) — constrained spacing/type/color/shadow scales, visual hierarchy through weight and color rather…

    595 GitHub stars~3.5k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Web Design

    drewnekota/cetus

    A skill your agent uses when building any web page, HTML artifact, landing page, dashboard, slide deck, report, email, or UI component the user will look at.

    146 GitHub stars~2.1k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • iOS Design Guidelines

    MacMagazine/app-iOS

    MacMagazine design reference — theme tokens, card selection, typography, touch targets, accessibility, dark mode.

    170 GitHub stars~767 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Atlas

    pacifio/ui

    Atlas — a design language and UI guide system for agent interfaces.

    154 GitHub stars~2.5k tokensUpdated 5 mo ago
    Frontend & DesignAuto-check passed
  • Frontend Design

    avibebuilder/claude-prime

    Builds distinctive, production-grade UIs that avoid generic AI aesthetics.

    120 GitHub stars~2.1k tokensUpdated 4 mo ago
    Frontend & DesignAuto-check passed

More from Jason904/ui-skill-lab

All 8 skills in this repo
  • Image2 Prompt Pack

    Jason904/ui-skill-lab

    This skill should be used when the user asks to convert a PRD, OpenSpec change, or product idea into a structured GPT Image 2 UI prompt pack for generating high-fidelity flat UI reference images.

    144 GitHub stars~1k tokensUpdated 3 mo ago
    Auto-check passed
  • UI UX Pro Max Skill

    Jason904/ui-skill-lab

    A skill your agent uses for UI/UX design intelligence in Codex: planning, building, reviewing, improving, or implementing web/mobile interfaces, landing pages, dashboards, design systems…

    144 GitHub stars~865 tokensUpdated 3 mo ago
    Auto-check passed
  • Design System Governor

    Jason904/ui-skill-lab

    This skill should be used after visual-spec-review passes and before OpenSpec UI update or frontend implementation.

    144 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Spec Compliance Review

    Jason904/ui-skill-lab

    This skill should be used after Codex implements a UI change and before visual acceptance review.

    144 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Visual Acceptance Review

    Jason904/ui-skill-lab

    This skill should be used after spec-compliance-review and before visual-diff-fix.

    144 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Visual Diff Fix

    Jason904/ui-skill-lab

    This skill should be used after visual-acceptance-review has generated fix-tasks.json.

    144 GitHub stars~978 tokensUpdated 3 mo ago
    Auto-check passed

Questions about Visual To Spec

What does Visual To Spec do?

A skill your agent uses when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation…. Visual To Spec is an agent skill from Jason904/ui-skill-lab. Use when a user provides a reference UI screenshot and asks to extract visual specs, generate design tokens, create a component tree, or produce a frontend implementation contract, including dashboards, landing pages, mobile screens, image-led UIs, panorama scenes, 3D exhibit UIs, and screenshot-to-spec tasks.

When should I use Visual To Spec?

Visual To Spec fits situations like: A user provides a reference UI screenshot and asks to extract visual specs; generate design tokens; create a component tree; produce a frontend implementation contract.

How do I install Visual To Spec in Claude Code?

Run `npx skills add Jason904/ui-skill-lab --skill visual-to-spec -a claude-code`. Or copy the skill folder (skills/visual-to-spec in Jason904/ui-skill-lab) into .claude/skills/visual-to-spec in your project. Claude Code loads it when a task matches its description.

How do I install Visual To Spec in Codex?

Run `npx skills add Jason904/ui-skill-lab --skill visual-to-spec -a codex`. Or copy the skill folder (skills/visual-to-spec in Jason904/ui-skill-lab) into .agents/skills/visual-to-spec in your project. Codex loads it when a task matches its description.

Can I use Visual To Spec in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jason904/ui-skill-lab --skill visual-to-spec -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/visual-to-spec, .gemini/skills/visual-to-spec, .github/skills/visual-to-spec and .opencode/skills/visual-to-spec in your project.

What does Visual To Spec need to run?

SKILL.md names no scripts, command-line tools or credentials: Visual To Spec is instructions for the agent only.

Does Visual To Spec access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Visual To Spec safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Visual To Spec use?

Visual To Spec is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Visual To Spec use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.

What are the alternatives to Visual To Spec?

Skills that share tags, products or a category with Visual To Spec: Web Design (xiaopu-ai/web-design, 783 stars), Refactoring UI (s0xDk/refactoring-ui-skill, 595 stars), Web Design (drewnekota/cetus, 146 stars) and iOS Design Guidelines (MacMagazine/app-iOS, 170 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Visual To Spec?

Jason904 (a GitHub user) maintains it in Jason904/ui-skill-lab, which has 144 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on June 29, 2026.

Source: Jason904/ui-skill-lab on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.