Agent skill

Design Critique

by WrongStack in WrongStack/WrongStack

A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…

MITAuto-check passedMedia & Creative

Install Design Critique

skills CLI
$ npx skills add WrongStack/WrongStack --skill design-critique -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WrongStack/WrongStack design-critique --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WrongStack/WrongStack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/skills/design-critique .claude/skills/design-critique && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
design-critique
GitHub stars
368
Token cost
~3k tokens
SKILL.md length
1,506 words
Files
2 (incl. references)
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…

  • Works in 5 steps: Ground truth → Machine pass → Craft pass → …
  • Audit an interface that already exists and say precisely why it looks generated
  • SKILL.md covers Overview, Workflow, Rules and Anti-patterns in critiques, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Design Critique is an agent skill from WrongStack/WrongStack. Use this skill to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states, accessibility and copy, ending in a ranked fix list. Triggers: user says "review the design", "critique this UI", "why does this look bad", "looks generic", "looks AI-generated", "design review", "audit the UI", "make this look professional", "what's wrong with this page", "design feedback".

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/rubric.md`).

It sits in Media & Creative, covering Design review and critique. The repository describes itself as: An AI coding agent that reads your code, edits files, runs commands, and reasons through bugs — across a terminal REPL, a full-screen TUI, and a browser UI, while you keep your… The licence is MIT.

When your agent uses it

  • Audit an interface that already exists and say precisely why it looks generated
  • Unfinished — a scored rubric across composition
  • Accessibility and copy
  • Ending in a ranked fix list

Example prompts

  • “review the design”
  • “critique this UI”
  • “why does this look bad”
  • “/design-critique”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Ground truth
  2. Machine pass
  3. Craft pass
  4. Score and rank
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit ec76a20. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Design Critique loads about 3k tokens when it runs, and up to ~4.6k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,506 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from WrongStack/WrongStack at commit ec76a20, republished under its MIT licence (© WrongStack). 1,506 words, ~2,952 tokens.

Download SKILL.mdSave it as .claude/skills/design-critique/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
design-critique
description
Use this skill to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states, accessibility and copy, ending in a ranked fix list. Triggers: user says "review the design", "critique this UI", "why does this look bad", "looks generic", "looks AI-generated", "design review", "audit the UI", "make this look professional", "what's wrong with this page", "design feedback".
version
2.0.0
required-capabilities
filesystem.read
required-tools
design, skill, read, grep
optional-capabilities
browser.interact, verification.run

Design Critique — WrongStack

Overview

A design critique that says "it looks a bit generic, maybe add more spacing" is worthless. This skill produces a scored, evidenced, ranked audit: each finding names the file and line, the rule it breaks, and the concrete replacement — the same standard a code review is held to.

Two independent failure classes, always reported separately:

  • Adherence — does the code use the kit's tokens? Machine-checkable.
  • Craft — does the result look designed? Needs judgment, guided by a rubric.

A UI at 100% adherence with a failing craft score is the common case, and saying so plainly is the whole value of this skill.

Workflow

1. Establish ground truth   → active kit + brief
2. Machine pass             → design {action:"verify"}
3. Craft pass               → rubric, six axes, evidence per finding
4. Score + rank             → what to fix first, what to ignore
5. Report                   → findings, not adjectives
1 — Ground truth
design {action:"list"}        # what is pinned

Read .design/brief.md with read if it exists — a critique that contradicts a decision the team already made is noise, unless the decision itself is the problem (say so explicitly, once). Read .design/rules.md for project overrides, and grep the UI source for token usage to ground the adherence findings.

Adherence means different things in three cases. Establish which one you are in before reporting any percentage.

CaseWhat adherence means here
A kit is pinnedRun the machine pass and report its score as-is.
No kit, but the project has its own design system — a theme file, semantic CSS variables, a token layerMeasure against the project's own tokens, not a kit. Find the token source (@theme block, :root variables, theme constants), then look for literals that bypass it. A percentage produced by pinning an arbitrary kit is meaningless — do not report one.
No kit and no token layerThat is finding #1: the UI has no source of truth, and every color is a decision nobody can revisit.

The middle case is the common one in a mature codebase, and mis-reporting it as the third is how a critique loses the room: the team already built a system, and being told they have none is simply wrong. Say instead which axis of their system is missing — a kit carries radius, spacing, type, motion and elevation, and a hand-rolled system is usually missing one of them.

2 — Machine pass
design {action:"verify"}

Check what the machine pass could actually see. It reads utility classes and CSS; a native-stack screen (react-native, flutter, swiftui, compose) has neither, so it returns a clean result on a file it never checked. If verify reports files with no class/utility signal, say so in the report — Tokens: not machine-checkable (native stack) — and carry the entire adherence judgement by reading the theme constants and spacing scale yourself. Never let a vacuous clean pass stand in as evidence.

Run design with {action:"verify"} and report the breakdown by axis. Treat composition findings as craft evidence, not token drift — they are patterns that are token-clean and still generic.

3 — Craft pass

Inspect a rendered screen or supplied image before making visual claims. Record the route/artifact, viewport, theme and state. Source alone supports implementation findings, not claims that a screen looks balanced, passes contrast or behaves correctly. Mark missing evidence unverified, not n/a; omit the overall score when an applicable axis lacks evidence. Do not invent screenshot observations.

Judge against the primary task and brief. Repeated rows, equal cards, symmetry, a single font or a gradient can be appropriate. A scanner match is a question to investigate, not a verdict; confirm the visible problem before requesting a change. Check whether a logo swap leaves an unrelated but equally plausible product, and whether the interface uses the actual domain's content and workflow.

Classify the surface first. Two axes are scored differently depending on it, and some are not scorable at all:

SurfaceExamplesWhat changes
Pagelanding, marketing, docs, article, onboardingNothing — every axis applies as written
App screendashboard, console, table view, settings, editor chromeStructure: "centered monotony" and "hero + three cards" do not apply; judge tile/lane rhythm and whether one element earns the focal point. Typography: the 60–75ch measure rule applies to prose blocks only, never to tables, labels or numeric cells
Kiosk / public terminalticket machine, self-checkout, wayfinding panel, check-in screenEvery axis applies, plus the physical checks below — which no other surface needs and which outrank taste when they conflict
Component in isolationone primitive, one cardStructure and Copy are usually not scorable

Kiosk is not a small page. Its constraints are physical, and a rubric that only asks design questions will hand a kiosk a flattering score for the wrong reasons. Ask these as well, and treat a failure as blocking rather than craft:

  • Readable at distance. Can the primary state and the total/answer be read from two metres, in glare? If the answer depends on leaning in, it fails.
  • One cold finger. Targets well above the 44px floor (64px+), spaced so a mis-tap cannot select the neighbour. No hover, no drag, no long-press, no dropdown.
  • Recoverable. Every destructive or committing step has a visible way back. A stranded user cannot refresh, log in again, or email support.
  • No session. Nothing personal persists on screen, and the screen returns to its start state on its own after inactivity.
  • Standing, not sitting. Content sits in the upper-middle band; nothing essential lives at the very bottom of a tall panel.

Score the six craft axes as usual, then report the physical checks as a separate pass/fail list. A kiosk that scores 4/5 on craft and fails "readable at distance" is a failing screen.

Never score an axis that does not apply to the surface. Write n/a (app screen) and move on — a fabricated 3/5 drags the overall score, which is the lowest axis, and makes the whole report meaningless.

Score each remaining axis 0–5 against the rubric, loaded with the skill tool:

skill({ name: "design-critique", resource: "references/rubric.md" })

Never score from feel — each score cites at least one concrete observation.

AxisThe question it answers
StructureIs there a grid and a focal point, or an even mat of equal blocks?
TypographyDoes hierarchy survive in greyscale? Is the measure controlled?
ColorDo semantic roles and visual emphasis support the task? Are supported themes deliberate and readable?
Surface & depthOne coherent elevation strategy, or borders+shadows stacked at random?
States & edgesEmpty, loading, error, overflow, long strings — present or happy-path only?
Copy & voiceDomain-specific, or interchangeable marketing filler?

Run the three cheap tests and report the result of each:

  • Greyscale — remove color. Hierarchy intact?
  • Squint — blur it. Is there one focal point per screenful?
  • Swap — replace the product name throughout. Does the copy still make sense for a different product? If yes, the copy is filler.
Show full SKILL.md (430 more words)Show less
4 — Score and rank

Overall = the lowest axis, not the average. One broken axis is what people see. Rank fixes by visible impact per unit of work; structure fixes almost always outrank color fixes.

Mark each finding:

  • Blocking — ships as broken (a11y failure, unreadable contrast, missing error state, horizontal scroll at 320px)
  • Craft — the taste gap; why it reads as generated
  • Nit — genuinely optional
5 — Report
## Design critique — <surface>

Surface: <page | app screen | component>
Tokens: <kit id · adherence pct> | <project's own system — no kit pinned> | <none>
Evidence: <route/artifact, viewport, theme, state; source-only gaps>
Craft: structure 2/5 · type 3/5 · color unverified · surface 3/5 · states 1/5 · copy 2/5
Verdict: <one sentence naming the single biggest reason it reads as generated>

### Blocking
1. `src/app/page.tsx:41` — no focus ring on the primary action …
### Craft
2. `src/app/page.tsx:12` — gradient-filled headline; hierarchy outsourced to a filter.
   Replace with: display size 2.75rem / weight 600, body dropped to muted.
### Nits
…

Rules

  1. Evidence or silence. Every finding names file:line or a named screen region. No "the spacing feels off".
  2. Each finding ships its replacement. Naming the flaw is half the work.
  3. Separate adherence from craft. Never let a clean verify imply the design is good, and never report a craft opinion as a token violation.
  4. Verdict first, in one sentence. The single biggest reason. If you cannot name one, you have not finished looking.
  5. Respect the requested outcome. For an audit-only request, report findings. When improvements are already authorized, record the evidence, implement the ranked fixes and inspect again without asking for the same authorization.
  6. Respect the brief. Decisions already made are not findings unless they are the problem; say that once, don't re-litigate.
  7. No praise padding. One line of what genuinely works, then the findings.
  8. Cap the list. Top 10 by impact; state the count of the remainder.
  9. Classify the surface and the token situation before scoring. Scoring a dense table view against page rules, or reporting an adherence percentage against a kit the project never adopted, produces confident nonsense.
  10. File each finding against the file that owns it. A screen with no focus ring may be importing a primitive that lacks one — the finding belongs to the primitive. Check before you blame the surface you happen to be reading.

Anti-patterns in critiques

Don'tDo
"Feels a bit generic""Three identical centered sections; no spine — evidence: lines 20, 48, 76"
"Add more whitespace""Section padding is uniform p-6; the kit's density scale gives 12/8/6 by role"
"Improve the colors""Secondary badges compete with the primary action in the inspected viewport; reduce their emphasis while preserving status meaning"
Scoring every axis 3/5Scores that differ, each with a citation
Mixing observations and intended fixesRecord the observed issue, implement authorized fixes, then report rechecked behavior

Skills in scope

  • design-craft — the rules this rubric audits against; use it to do the fixing
  • design-system — kit, tokens, and the machine adherence pass
  • web-platform-baseline — before claiming a capability is unavailable
  • code-review — when the findings are structural code problems, not design
  • output-standards — <nextsteps> shape when handing the fix list back

© WrongStack, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in packages/core/skills/design-critique of WrongStack/WrongStack.

  • SKILL.md
  • references/rubric.md

Open the folder on GitHubat commit ec76a20

Compare with similar skills

Design Critique next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Design Critique compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Design Critique this skillWrongStack/WrongStack368—~3kAutomated safety check: PassMIT
Reimagine It AuditKayforkind/reimagine-it203—~756Automated safety check: PassMIT
Design Reviewtsubotax/melta-ui201—~504Automated safety check: PassMIT
Canva Design Feedbackcanva-sdks/canva-skills115—~1.2kAutomated safety check: PassApache-2.0
Design Critiquepaperclipai/paperclip98k—~1.2kAutomated safety check: PassMIT
Steve Jobs Design Reviewwondelai/skills2.4k—~4.5kAutomated safety check: PassMIT

Similar skills

  • Reimagine It Audit

    Kayforkind/reimagine-it

    Design Health — runs 19 deterministic quality checks on HTML output.

    203 GitHub stars~756 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Design Review

    tsubotax/melta-ui

    HTMLファイルをmelta UIデザインシステムに照らしてレビューし、違反を検出・分類・修正提案する。トリガー: 「デザインレビュー」「DSチェック」「禁止パターンチェック」「design review」「check compliance」「DS準拠確認」。対象ファイルのパスを引数で受け取る。

    201 GitHub stars~504 tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Canva Design Feedback

    canva-sdks/canva-skills

    Read a Canva design and return structured, actionable design feedback — visual hierarchy, copy/messaging, layout & spacing, consistency, readability, and accessibility.

    115 GitHub stars~1.2k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Design Critique

    paperclipai/paperclip

    Give a structured product design critique — user job clarity, hierarchy, affordance, error states, accessibility, and consistency — focused on what to change, in what order, and why.

    98k GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Review designs, products, and features with Steve Jobs' standards: ruthless simplicity, focus, and end-to-end excellence.

    2.4k GitHub stars~4.5k tokensUpdated 26 days ago
    Media & CreativeAuto-check passed
  • Critique

    sanqiufong/slides-from-anything

    Run a 5-dimension expert design review on any HTML artifact in the project — Philosophy / Visual hierarchy / Detail / Functionality / Innovation, each scored 0–10.

    132 GitHub starsUsed in 1 repo~2.4k tokens
    Media & CreativeAuto-check passed

More from WrongStack/WrongStack

All 38 skills in this repo
  • Design Craft

    WrongStack/WrongStack

    Design or substantially improve user-facing interfaces with a product-specific visual direction, content hierarchy, and rendered critique.

    368 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Mailbox Bridge

    WrongStack/WrongStack

    A skill your agent uses when external coding agents (Claude Code, Aider, custom scripts) need to participate in the project's shared WrongStack mailbox, or when a user asks to "expose the mailbox"…

    368 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Multi Agent

    WrongStack/WrongStack

    A skill your agent uses whenever work can be split across multiple AI agents running in parallel, or when orchestrating leader/worker patterns in WrongStack.

    368 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Web Platform Baseline

    WrongStack/WrongStack

    Use this skill before asserting that a CSS, HTML or accessibility capability is available, unavailable, or the right tool — it carries dated, refreshable platform facts and refuses to let stale…

    368 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Wrongstack Mailbox

    WrongStack/WrongStack

    A skill your agent uses when the user wants to communicate with WrongStack's shared project mailbox from outside WrongStack — read messages sent by WrongStack agents, send replies, broadcast to all…

    368 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • API Design

    WrongStack/WrongStack

    A skill your agent uses when designing, implementing, or reviewing an HTTP API — endpoints, request and response shapes, errors, pagination, versioning, and authorization.

    368 GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Questions about Design Critique

What does Design Critique do?

A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…. Design Critique is an agent skill from WrongStack/WrongStack. Use this skill to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states, accessibility and copy, ending in a ranked fix list.

When should I use Design Critique?

Design Critique fits situations like: audit an interface that already exists and say precisely why it looks generated; unfinished — a scored rubric across composition; accessibility and copy; ending in a ranked fix list.

How do I install Design Critique in Claude Code?

Run `npx skills add WrongStack/WrongStack --skill design-critique -a claude-code`. Or copy the skill folder (packages/core/skills/design-critique in WrongStack/WrongStack) into .claude/skills/design-critique in your project. Claude Code loads it when a task matches its description.

How do I install Design Critique in Codex?

Run `npx skills add WrongStack/WrongStack --skill design-critique -a codex`. Or copy the skill folder (packages/core/skills/design-critique in WrongStack/WrongStack) into .agents/skills/design-critique in your project. Codex loads it when a task matches its description.

Can I use Design Critique in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WrongStack/WrongStack --skill design-critique -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-critique, .gemini/skills/design-critique, .github/skills/design-critique and .opencode/skills/design-critique in your project.

What does Design Critique need to run?

SKILL.md names no scripts, command-line tools or credentials: Design Critique is instructions for the agent only.

Does Design Critique access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Design Critique safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Design Critique use?

Design Critique is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Design Critique use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Design Critique?

Skills that share tags, products or a category with Design Critique: Reimagine It Audit (Kayforkind/reimagine-it, 203 stars), Design Review (tsubotax/melta-ui, 201 stars), Canva Design Feedback (canva-sdks/canva-skills, 115 stars) and Design Critique (paperclipai/paperclip, 98k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Design Critique?

WrongStack (a GitHub organization) maintains it in WrongStack/WrongStack, which has 368 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 6, 2026.

Source: WrongStack/WrongStack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.