Agent skill

Design Evaluation

by SeanJ1ang in SeanJ1ang/design-judge-skills

Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.

Apache-2.0Auto-check passedEducation

Install Design Evaluation

skills CLI
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/design-evaluation .claude/skills/design-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
design-evaluation
GitHub stars
712
Token cost
~3.1k tokens
SKILL.md length
1,364 words
Files
59 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.

  • Works in 8 steps: Confirm the user-selected maturity → Build the evaluation profile → Build the evidence ledger → …
  • A user asks to judge
  • SKILL.md covers Purpose, Scope Boundary, Required User Input and Evaluation Workflow, plus 2 more sections
  • Rank designs by evidence-aligned evaluation score

What it does

Design Evaluation is an agent skill from SeanJ1ang/design-judge-skills. Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric. Classify each work, score design quality and presentation, identify Critical risks, report evidence confidence, and optionally shortlist works within separate maturity tracks. Use when a user asks to judge, score, critique, review, diagnose, batch-evaluate, or rank designs by evidence-aligned evaluation score. Do not use this skill to retrieve winners, choose an award, produce a redesign, audit…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 63 other files, including scripts and reference files (for example `README.md`, `README_EN.md` and `agents/openai.yaml`).

It sits in Education, covering Quizzes and assessments and Design review and critique. The repository describes itself as: Evidence-driven Agent Skills for design award research, evaluation, award matching, entry writing, and submission readiness. The licence is Apache-2.0.

When your agent uses it

  • A user asks to judge
  • Rank designs by evidence-aligned evaluation score
  • Retrieve winners
  • Choose an award

Example prompts

  • “/design-evaluation”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Confirm the user-selected maturity
  2. Build the evaluation profile
  3. Build the evidence ledger
  4. Load the rubric
  5. Score from evidence
  6. Identify findings
  7. Apply optional benchmark evidence
  8. Present the result

What it can do on your machine

Read from SKILL.md and the folder at commit abf53e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Design Evaluation loads about 3.1k tokens when it runs, and up to ~146k if it reads all its reference files. Until then it costs about 152 tokens; SKILL.md has 1,364 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~152
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~146k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from SeanJ1ang/design-judge-skills at commit abf53e6, republished under its Apache-2.0 licence (© SeanJ1ang). 1,364 words, ~3,133 tokens.

Download SKILL.mdSave it as .claude/skills/design-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.
name
design-evaluation
description
Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric. Classify each work, score design quality and presentation, identify Critical risks, report evidence confidence, and optionally shortlist works within separate maturity tracks. Use when a user asks to judge, score, critique, review, diagnose, batch-evaluate, or rank designs by evidence-aligned evaluation score. Do not use this skill to retrieve winners, choose an award, produce a redesign, audit submission-file compliance, simulate an official jury, or predict winning probability.

Design Evaluation

Purpose

Evaluate design quality consistently without pretending that a score is an award outcome. Keep design quality, presentation quality, and evidence confidence separate. Require the user to choose the maturity track.

Scope Boundary

  • Evaluate the supplied design and supplied presentation materials.
  • Classify one primary discipline, one primary sector, and optional secondary labels and focus tags.
  • Build an evidence ledger before scoring.
  • Score the general rubric and report Critical findings separately.
  • Apply an optional award-aligned lens only when the user already names a target award.
  • Batch-evaluate a fixed corpus only after the user approves the maturity mapping for every included record.
  • Produce score-based shortlists within each maturity track; never cross-rank Student Concept and Mature Work.

Do not:

  • infer or change the work maturity;
  • retrieve award winners inside this skill;
  • recommend which award to enter;
  • turn findings into a full redesign proposal;
  • audit upload limits, filenames, declarations, licences, or portal compliance;
  • call the result an official iF, Red Dot, or other jury decision;
  • estimate an exact probability of winning.
  • label a rank percentile, top-decile membership, or score as a winning probability.

Route winner retrieval to $design-award-search, award selection to $design-award-match, concrete redesign work to $design-optimization when available, and final package compliance to $design-submission-check.

Required User Input

Accept images, a PDF, project text, a portfolio page, video frames, prototype evidence, test records, or a structured brief.

Maturity is mandatory and must come from the user. Accept exactly:

  • Student Concept / 学生概念
  • Mature Work / 成熟作品

If maturity is absent, ask exactly one question and stop scoring:

请选择作品成熟度:“学生概念”或“成熟作品”。

Never infer maturity from the author's identity, image finish, prototype appearance, commercial branding, or supplied metadata. If evidence conflicts with the selected maturity, preserve the user's selection and record Maturity evidence mismatch.

For a batch, an explicit user-approved mapping rule counts as user selection for every record matched by that rule. Reject unmatched values rather than inferring them. Record the mapping rule and maturity_source: user in the batch manifest.

Offer this template when the user asks how to use the skill:

text
Project: {name}
Maturity: Student Concept | Mature Work  # selected by the user
Primary function: {what it does}
Target user: {who uses it}
Use context: {where and when}
Materials: {attachments or links}
Evaluation mode: General | optional named award-aligned lens

Evaluation Workflow

For batch work, first read references/batch-evaluation.md. Use scripts/batch_evaluation.py for deterministic scoring, failure isolation, and separate-track shortlisting. Use a project adapter for private database access; never bundle database rows, images, signed URLs, or credentials in the public Skill.

1. Confirm the user-selected maturity

Record:

yaml
maturity: student_concept | mature_work
maturity_source: user

Do not proceed with a numeric score when maturity_source is missing or is not user.

2. Build the evaluation profile

Read references/classification-policy.md and references/profiles/classification.json.

Extract:

  • primary function, target user, use context, and claimed outcome;
  • one primary design discipline and up to two secondary disciplines;
  • one primary application sector and up to one secondary sector;
  • zero or more focus tags;
  • supplied material types and obvious material limitations.

The classification confidence is separate from evaluation confidence. Ask no additional question when a reasonable classification can be stated as an assumption.

3. Build the evidence ledger

Read references/evidence-policy.md. For every scored dimension, assign exactly one evidence state:

  • Verified
  • Supported
  • Claimed
  • Missing

Attach concise evidence references and distinguish observable facts from author claims and evaluator inference.

4. Load the rubric

Read references/evaluation-framework.md.

Load:

  1. references/profiles/core.json;
  2. the user-selected maturity profile;
  3. the relevant classification overlay in references/profiles/sector-overlays.json;
  4. an optional aggregate benchmark context resolved by scripts/benchmark_profiles.py;
  5. an optional file from references/profiles/award-lenses/ when the user names that target.

Award lenses produce a separate alignment section. Never replace or mathematically blend the general score with an award-aligned result.

For the main iF context, read references/if-benchmark-methodology.md. Resolve an exact normalized category profile first, then its mapped discipline profile, then the core fallback.

For iF Student context, read references/if-student-benchmark-methodology.md. Load it only after the user has selected student_concept. Reject it for mature_work. Treat its 15 SDG categories as issue themes, never as evidence of product, communication, interface, spatial, or other design discipline. Resolve a high-sample SDG theme first and otherwise use the competition-wide student profile.

For Red Dot context, read references/red-dot-benchmark-methodology.md. Keep Product Design, Brands & Communication Design, and Design Concept separate. Resolve an exact high-sample category first, then an explicitly supplied competition line, then the mapped evaluation discipline, then the core fallback. Never infer the Red Dot competition line from maturity.

For IDEA context, read references/idea-benchmark-methodology.md. Resolve an exact high-sample category first, then a supplied or provisionally mapped discipline, then the program-wide context. Treat its discipline profiles as multi-label and non-additive. Load Student Designs only for student_concept, and never infer discipline from that category.

Treat every bundled benchmark profile as observed winner context with score_effect: none; use it to choose evidence questions and explain presentation coverage, never to change weights or predict an award result.

5. Score from evidence

Assign a raw score from 0 to 5 to all seven design dimensions and six presentation dimensions. Give one short reason for every score.

The general score allocates 50 points to design quality and 50 points to presentation quality.

Use scripts/score_evaluation.py for deterministic weighting and profile validation. The score is:

text
weighted contribution = raw score / 5 * dimension weight
general score = design score + presentation score

Apply evidence caps from the framework. Evidence confidence does not otherwise add or subtract arbitrary points.

Show full SKILL.md (536 more words)Show less
6. Identify findings

Separate:

  • Critical: a fundamental contradiction, unverified safety-critical mechanism, serious harm risk, or a failure that invalidates a core claim;
  • Major: materially reduces design quality or credibility but does not invalidate the whole proposal;
  • Minor: local weakness with limited effect.

Critical findings cannot be cancelled by the average score. Evaluation findings describe the problem and its consequence; leave detailed redesign instructions to the optimization module.

7. Apply optional benchmark evidence

Consume either:

  • a verified benchmark set supplied by the user or produced by $design-award-search; or
  • the bundled aggregate iF context when the user targets iF or an iF category can be resolved; or
  • the bundled aggregate iF Student context only when the user-selected maturity is student_concept.
  • the bundled aggregate Red Dot context when Red Dot is the named target or its category and competition line are supplied.
  • the bundled aggregate IDEA context when IDEA is the named target or an IDEA category is supplied.

For bundled iF context, run:

text
python scripts/benchmark_profiles.py --category "{raw or normalized iF category}" --discipline "{evaluation discipline}" --pretty

For iF Student, run:

text
python scripts/benchmark_profiles.py --source if_student_observed_winners --maturity student_concept --category "{raw or normalized SDG theme}" --pretty

For Red Dot, run:

text
python scripts/benchmark_profiles.py --source red_dot_observed_winners --category "{raw or normalized Red Dot category}" --competition "{Product Design | Brands & Communication Design | Design Concept}" --discipline "{evaluation discipline}" --pretty

For IDEA, run:

text
python scripts/benchmark_profiles.py --source idea_observed_recognized --maturity "{student_concept | mature_work}" --category "{raw or normalized IDEA category}" --discipline "{evaluation discipline}" --pretty

Do not run or load the iF Student benchmark for mature_work; the resolver must raise an error. Do not use an SDG theme to infer or replace the user's project discipline. For Red Dot, treat the competition line as explicit target context and never infer it from the project maturity. For IDEA, reject Student Designs when maturity is mature_work; its student category must not infer or replace project discipline.

State the matched profile, fallback used, sample size, years, review status, and limitations. Keep the benchmark section outside the numeric score calculation. Use it to explain differentiation, missing evidence, and presentation coverage, not hidden jury preferences or winning probability.

8. Present the result

Read references/output-template.md. Lead with:

  • user-selected maturity;
  • classification;
  • overall score, design score, presentation score;
  • evidence confidence;
  • Critical status.

Include the complete dimension table and evidence gaps. End with an explicit limitation statement.

Scoring Guardrails

  • Student Concept and Mature Work are separate tracks; do not rank them directly by total score.
  • iF Student benchmark evidence is valid only in the Student Concept track.
  • Do not reward unsupported claims because they sound plausible.
  • Do not convert missing evidence into invented defects.
  • Do not score an inaccessible or unreadable visual detail as observed.
  • Keep category relevance and material coverage explicit.
  • A high score means strong evidence-aligned design quality under this rubric, not an award forecast.
  • Use the term {award}-aligned assessment, never {award} official jury simulation.
  • Call batch output evidence-aligned evaluation shortlist or within-track top decile, never award-probability shortlist.
  • Require all expected records to reach a terminal status before finalizing a corpus-wide shortlist; report failed and non-evaluable records separately.
  • Exclude unresolved Critical findings and Low-confidence evaluations from the primary shortlist. Preserve them in review queues rather than deleting them.
  • Resolve exact ties at the cutoff by including every work with the same total score; report when this makes the shortlist larger than the nominal ratio.

Example Invocations

  • Use $design-evaluation on this fixed database snapshot. Apply my confirmed maturity mapping, score every evaluable work, and return the within-track top 10% evidence-aligned shortlist. Do not report award probability.

  • 使用 $design-evaluation 评价附件中的学生概念。成熟度由我确定为“学生概念”。

  • Use $design-evaluation. Maturity: Mature Work. Evaluate the product photos, test summary, and entry boards.

  • 使用 $design-evaluation,以通用体系评价,并附加 iF-aligned 维度映射;不要预测获奖概率。

© SeanJ1ang, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 58 other files (scripts, references) in skills/design-evaluation of SeanJ1ang/design-judge-skills.

  • SKILL.md
  • README.md
  • README_EN.md
  • agents/openai.yaml
  • examples/evaluation-input.example.json
  • references/batch-evaluation.md
  • references/classification-policy.md
  • references/design-evaluations-schema.sql
  • references/evaluation-framework.md
  • references/evidence-policy.md
  • references/idea-benchmark-methodology.md
  • references/if-benchmark-methodology.md
  • references/if-student-benchmark-methodology.md
  • references/output-template.md
  • references/profiles/award-lenses/if-design.json
  • references/profiles/award-lenses/if-student.json
  • … and 43 more

Open the folder on GitHubat commit abf53e6

Compare with similar skills

Design Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Design Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Design Evaluation this skillSeanJ1ang/design-judge-skills712—~3.1kAutomated safety check: PassApache-2.0
Skill JudgeshareAI-lab/lab-skills314—~1.9kAutomated safety check: PassApache-2.0
Skill Reviewerdaymade/claude-code-skills1.4k—~2.1kAutomated safety check: PassMIT
Exceptional Web Designmicrosoft/power-platform-skills979—~3.1kAutomated safety check: NotesMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT

Similar skills

  • Skill Judge

    shareAI-lab/lab-skills

    Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

    314 GitHub stars~1.9k tokensUpdated 23 days ago
    EducationAuto-check passed
  • Skill Reviewer

    daymade/claude-code-skills

    Reviews skill quality with evidence-based design rubrics and read-only batch inventories.

    1.4k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Exceptional Web Design

    microsoft/power-platform-skills

    Official

    Reviews the design of an existing Power Pages site - a live URL or a local project folder - and recommends what to change, without modifying anything.

    979 GitHub stars~3.1k tokensUpdated today
    Frontend & DesignAuto-check: notes
  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 2 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed

More from SeanJ1ang/design-judge-skills

  • Design Award Match

    SeanJ1ang/design-judge-skills

    Match a design project to supported design-award programs, tracks, and entry categories; apply structural eligibility gates; verify current official rules; compare published criteria and cautiously…

    712 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Design Award Search

    SeanJ1ang/design-judge-skills

    Find and verify award-winning designs in the same or adjacent functional category through eight explicit relevance dimensions: problem and user, core function, sensing technology, intervention…

    712 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Design Information Prep

    SeanJ1ang/design-judge-skills

    Extract evidence-grounded project facts from user-provided design attachments, identify missing information, and prepare the exact written fields required by supported design-award entry forms.

    712 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Design Submission Check

    SeanJ1ang/design-judge-skills

    Audit a design-award submission package against the current official rules for a specific award cycle.

    712 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Design Award Pipeline

    SeanJ1ang/design-judge-skills

    Route and coordinate an end-to-end design-award workflow across winner research, evidence-based evaluation, award matching, entry-text preparation, and final submission checking.

    712 GitHub stars~725 tokensUpdated 1 mo ago
    Auto-check passed
  • Design Judge Shared

    SeanJ1ang/design-judge-skills

    Shared support package for the Design Judge skill collection.

    712 GitHub stars~176 tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Design Evaluation

What does Design Evaluation do?

Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric. Design Evaluation is an agent skill from SeanJ1ang/design-judge-skills. Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.

When should I use Design Evaluation?

Design Evaluation fits situations like: A user asks to judge; rank designs by evidence-aligned evaluation score; retrieve winners; choose an award.

How do I install Design Evaluation in Claude Code?

Run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a claude-code`. Or copy the skill folder (skills/design-evaluation in SeanJ1ang/design-judge-skills) into .claude/skills/design-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Design Evaluation in Codex?

Run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a codex`. Or copy the skill folder (skills/design-evaluation in SeanJ1ang/design-judge-skills) into .agents/skills/design-evaluation in your project. Codex loads it when a task matches its description.

Can I use Design Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-evaluation, .gemini/skills/design-evaluation, .github/skills/design-evaluation and .opencode/skills/design-evaluation in your project.

What does Design Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Design Evaluation is instructions for the agent only. Our summary lists: Python 3.

Does Design Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Design Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Design Evaluation use?

Design Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Design Evaluation use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 143k tokens, read only when the agent opens those files.

What are the alternatives to Design Evaluation?

Skills that share tags, products or a category with Design Evaluation: Skill Judge (shareAI-lab/lab-skills, 314 stars), Skill Reviewer (daymade/claude-code-skills, 1.4k stars), Exceptional Web Design (microsoft/power-platform-skills, 979 stars) and DeepTutor CLI (HKUDS/DeepTutor, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Design Evaluation?

SeanJ1ang (a GitHub user) maintains it in SeanJ1ang/design-judge-skills, which has 712 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 24, 2026.

Source: SeanJ1ang/design-judge-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.