Skill Judge
shareAI-lab/lab-skills
Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.
Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/design-evaluation .claude/skills/design-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .claude/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/design-evaluation .agents/skills/design-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .agents/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/design-evaluation .cursor/skills/design-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .cursor/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/SeanJ1ang/design-judge-skills.git --path skills/design-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/design-evaluation .gemini/skills/design-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .gemini/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/design-evaluation .github/skills/design-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .github/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install SeanJ1ang/design-judge-skills design-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SeanJ1ang/design-judge-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/design-evaluation .opencode/skills/design-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "design-evaluation" agent skill from https://github.com/SeanJ1ang/design-judge-skills/tree/main/skills/design-evaluation into .opencode/skills/design-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
design-evaluationEvaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.
Design Evaluation is an agent skill from SeanJ1ang/design-judge-skills. Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric. Classify each work, score design quality and presentation, identify Critical risks, report evidence confidence, and optionally shortlist works within separate maturity tracks. Use when a user asks to judge, score, critique, review, diagnose, batch-evaluate, or rank designs by evidence-aligned evaluation score. Do not use this skill to retrieve winners, choose an award, produce a redesign, audit…
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 63 other files, including scripts and reference files (for example `README.md`, `README_EN.md` and `agents/openai.yaml`).
It sits in Education, covering Quizzes and assessments and Design review and critique. The repository describes itself as: Evidence-driven Agent Skills for design award research, evaluation, award matching, entry writing, and submission readiness. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit abf53e6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Design Evaluation loads about 3.1k tokens when it runs, and up to ~146k if it reads all its reference files. Until then it costs about 152 tokens; SKILL.md has 1,364 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from SeanJ1ang/design-judge-skills at commit abf53e6, republished under its Apache-2.0 licence (© SeanJ1ang). 1,364 words, ~3,133 tokens.
.claude/skills/design-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.Evaluate design quality consistently without pretending that a score is an award outcome. Keep design quality, presentation quality, and evidence confidence separate. Require the user to choose the maturity track.
Do not:
Route winner retrieval to $design-award-search, award selection to $design-award-match, concrete redesign work to $design-optimization when available, and final package compliance to $design-submission-check.
Accept images, a PDF, project text, a portfolio page, video frames, prototype evidence, test records, or a structured brief.
Maturity is mandatory and must come from the user. Accept exactly:
Student Concept / 学生概念Mature Work / 成熟作品If maturity is absent, ask exactly one question and stop scoring:
请选择作品成熟度:“学生概念”或“成熟作品”。
Never infer maturity from the author's identity, image finish, prototype appearance, commercial branding, or supplied metadata. If evidence conflicts with the selected maturity, preserve the user's selection and record Maturity evidence mismatch.
For a batch, an explicit user-approved mapping rule counts as user selection for every record matched by that rule. Reject unmatched values rather than inferring them. Record the mapping rule and maturity_source: user in the batch manifest.
Offer this template when the user asks how to use the skill:
Project: {name}
Maturity: Student Concept | Mature Work # selected by the user
Primary function: {what it does}
Target user: {who uses it}
Use context: {where and when}
Materials: {attachments or links}
Evaluation mode: General | optional named award-aligned lensFor batch work, first read references/batch-evaluation.md. Use scripts/batch_evaluation.py for deterministic scoring, failure isolation, and separate-track shortlisting. Use a project adapter for private database access; never bundle database rows, images, signed URLs, or credentials in the public Skill.
Record:
maturity: student_concept | mature_work
maturity_source: userDo not proceed with a numeric score when maturity_source is missing or is not user.
Read references/classification-policy.md and references/profiles/classification.json.
Extract:
The classification confidence is separate from evaluation confidence. Ask no additional question when a reasonable classification can be stated as an assumption.
Read references/evidence-policy.md. For every scored dimension, assign exactly one evidence state:
VerifiedSupportedClaimedMissingAttach concise evidence references and distinguish observable facts from author claims and evaluator inference.
Read references/evaluation-framework.md.
Load:
references/profiles/core.json;references/profiles/sector-overlays.json;scripts/benchmark_profiles.py;references/profiles/award-lenses/ when the user names that target.Award lenses produce a separate alignment section. Never replace or mathematically blend the general score with an award-aligned result.
For the main iF context, read references/if-benchmark-methodology.md. Resolve an exact normalized category profile first, then its mapped discipline profile, then the core fallback.
For iF Student context, read references/if-student-benchmark-methodology.md. Load it only after the user has selected student_concept. Reject it for mature_work. Treat its 15 SDG categories as issue themes, never as evidence of product, communication, interface, spatial, or other design discipline. Resolve a high-sample SDG theme first and otherwise use the competition-wide student profile.
For Red Dot context, read references/red-dot-benchmark-methodology.md. Keep Product Design, Brands & Communication Design, and Design Concept separate. Resolve an exact high-sample category first, then an explicitly supplied competition line, then the mapped evaluation discipline, then the core fallback. Never infer the Red Dot competition line from maturity.
For IDEA context, read references/idea-benchmark-methodology.md. Resolve an exact high-sample category first, then a supplied or provisionally mapped discipline, then the program-wide context. Treat its discipline profiles as multi-label and non-additive. Load Student Designs only for student_concept, and never infer discipline from that category.
Treat every bundled benchmark profile as observed winner context with score_effect: none; use it to choose evidence questions and explain presentation coverage, never to change weights or predict an award result.
Assign a raw score from 0 to 5 to all seven design dimensions and six presentation dimensions. Give one short reason for every score.
The general score allocates 50 points to design quality and 50 points to presentation quality.
Use scripts/score_evaluation.py for deterministic weighting and profile validation. The score is:
weighted contribution = raw score / 5 * dimension weight
general score = design score + presentation scoreApply evidence caps from the framework. Evidence confidence does not otherwise add or subtract arbitrary points.
Separate:
Critical: a fundamental contradiction, unverified safety-critical mechanism, serious harm risk, or a failure that invalidates a core claim;Major: materially reduces design quality or credibility but does not invalidate the whole proposal;Minor: local weakness with limited effect.Critical findings cannot be cancelled by the average score. Evaluation findings describe the problem and its consequence; leave detailed redesign instructions to the optimization module.
Consume either:
$design-award-search; orstudent_concept.For bundled iF context, run:
python scripts/benchmark_profiles.py --category "{raw or normalized iF category}" --discipline "{evaluation discipline}" --prettyFor iF Student, run:
python scripts/benchmark_profiles.py --source if_student_observed_winners --maturity student_concept --category "{raw or normalized SDG theme}" --prettyFor Red Dot, run:
python scripts/benchmark_profiles.py --source red_dot_observed_winners --category "{raw or normalized Red Dot category}" --competition "{Product Design | Brands & Communication Design | Design Concept}" --discipline "{evaluation discipline}" --prettyFor IDEA, run:
python scripts/benchmark_profiles.py --source idea_observed_recognized --maturity "{student_concept | mature_work}" --category "{raw or normalized IDEA category}" --discipline "{evaluation discipline}" --prettyDo not run or load the iF Student benchmark for mature_work; the resolver must raise an error. Do not use an SDG theme to infer or replace the user's project discipline. For Red Dot, treat the competition line as explicit target context and never infer it from the project maturity. For IDEA, reject Student Designs when maturity is mature_work; its student category must not infer or replace project discipline.
State the matched profile, fallback used, sample size, years, review status, and limitations. Keep the benchmark section outside the numeric score calculation. Use it to explain differentiation, missing evidence, and presentation coverage, not hidden jury preferences or winning probability.
Read references/output-template.md. Lead with:
Include the complete dimension table and evidence gaps. End with an explicit limitation statement.
{award}-aligned assessment, never {award} official jury simulation.evidence-aligned evaluation shortlist or within-track top decile, never award-probability shortlist.Use $design-evaluation on this fixed database snapshot. Apply my confirmed maturity mapping, score every evaluable work, and return the within-track top 10% evidence-aligned shortlist. Do not report award probability.
使用 $design-evaluation 评价附件中的学生概念。成熟度由我确定为“学生概念”。
Use $design-evaluation. Maturity: Mature Work. Evaluate the product photos, test summary, and entry boards.
使用 $design-evaluation,以通用体系评价,并附加 iF-aligned 维度映射;不要预测获奖概率。
© SeanJ1ang, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 58 other files (scripts, references) in skills/design-evaluation of SeanJ1ang/design-judge-skills.
Open the folder on GitHubat commit abf53e6
Design Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Design Evaluation this skillSeanJ1ang/design-judge-skills | 712 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Skill JudgeshareAI-lab/lab-skills | 314 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Skill Reviewerdaymade/claude-code-skills | 1.4k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Exceptional Web Designmicrosoft/power-platform-skills | 979 | — | ~3.1k | Automated safety check: Notes | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2k | Automated safety check: Pass | MIT |
shareAI-lab/lab-skills
Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.
daymade/claude-code-skills
Reviews skill quality with evidence-based design rubrics and read-only batch inventories.
microsoft/power-platform-skills
Reviews the design of an existing Power Pages site - a live URL or a local project folder - and recommends what to change, without modifying anything.
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
zarazhangrui/codebase-to-course
Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.
SeanJ1ang/design-judge-skills
Match a design project to supported design-award programs, tracks, and entry categories; apply structural eligibility gates; verify current official rules; compare published criteria and cautiously…
SeanJ1ang/design-judge-skills
Find and verify award-winning designs in the same or adjacent functional category through eight explicit relevance dimensions: problem and user, core function, sensing technology, intervention…
SeanJ1ang/design-judge-skills
Extract evidence-grounded project facts from user-provided design attachments, identify missing information, and prepare the exact written fields required by supported design-award entry forms.
SeanJ1ang/design-judge-skills
Audit a design-award submission package against the current official rules for a specific award cycle.
SeanJ1ang/design-judge-skills
Route and coordinate an end-to-end design-award workflow across winner research, evidence-based evaluation, award matching, entry-text preparation, and final submission checking.
SeanJ1ang/design-judge-skills
Shared support package for the Design Judge skill collection.
Categories
Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric. Design Evaluation is an agent skill from SeanJ1ang/design-judge-skills. Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.
Design Evaluation fits situations like: A user asks to judge; rank designs by evidence-aligned evaluation score; retrieve winners; choose an award.
Run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a claude-code`. Or copy the skill folder (skills/design-evaluation in SeanJ1ang/design-judge-skills) into .claude/skills/design-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a codex`. Or copy the skill folder (skills/design-evaluation in SeanJ1ang/design-judge-skills) into .agents/skills/design-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SeanJ1ang/design-judge-skills --skill design-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-evaluation, .gemini/skills/design-evaluation, .github/skills/design-evaluation and .opencode/skills/design-evaluation in your project.
SKILL.md names no scripts, command-line tools or credentials: Design Evaluation is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Design Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 143k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Design Evaluation: Skill Judge (shareAI-lab/lab-skills, 314 stars), Skill Reviewer (daymade/claude-code-skills, 1.4k stars), Exceptional Web Design (microsoft/power-platform-skills, 979 stars) and DeepTutor CLI (HKUDS/DeepTutor, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
SeanJ1ang (a GitHub user) maintains it in SeanJ1ang/design-judge-skills, which has 712 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 24, 2026.
Source: SeanJ1ang/design-judge-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.