Web Interface Guidelines Reviewer
vercel-labs/openreview
Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my…
Use after completing any non-trivial task. An agent skill from affaan-m/ECC.
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install affaan-m/ECC agent-self-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-self-evaluation .claude/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .claude/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install affaan-m/ECC agent-self-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agent-self-evaluation .agents/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .agents/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install affaan-m/ECC agent-self-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agent-self-evaluation .cursor/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .cursor/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/affaan-m/ECC.git --path skills/agent-self-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install affaan-m/ECC agent-self-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agent-self-evaluation .gemini/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .gemini/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install affaan-m/ECC agent-self-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agent-self-evaluation .github/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .github/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install affaan-m/ECC agent-self-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agent-self-evaluation .opencode/skills/agent-self-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-self-evaluation" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/agent-self-evaluation into .opencode/skills/agent-self-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-self-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-self-evaluationUse after completing any non-trivial task. An agent skill from affaan-m/ECC.
Agent Self Evaluation is an agent skill from affaan-m/ECC. Use after completing any non-trivial task. The agent self-rates its output on 5 axes — accuracy, completeness, clarity, actionability, conciseness — with concrete evidence per criterion. Produces a structured 1-5 scorecard with specific improvement suggestions.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `examples/high-score-example.md`, `examples/low-score-example.md` and `references/evaluation-criteria.md`).
It sits in Frontend & Design, covering Accessibility. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Self Evaluation loads about 1.9k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 665 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 665 words, ~1,890 tokens.
.claude/skills/agent-self-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.After completing a complex task, the agent pauses to rate its own output against a structured 5-axis rubric. This is NOT a pass/fail gate — it's a deliberate reflection step that catches omissions, flags overconfidence, and surface areas for improvement before the user has to.
references/hook-integration.md)| Axis | Question | What it catches |
|---|---|---|
| Accuracy | Are the facts, claims, and outputs correct? | Hallucinations, wrong API names, incorrect syntax, false statements |
| Completeness | Did it cover everything the user asked for? | Missed edge cases, unhandled error paths, forgotten requirements, skipped subtasks |
| Clarity | Is the explanation understandable and well-structured? | Confusing explanations, jargon without definition, missing context, rambling |
| Actionability | Can the user act on the output immediately? | Vague suggestions, missing steps, "you should X" without showing how, no verification path |
| Conciseness | Did it use the minimum words/tokens needed? | Redundancy, over-explanation, repeating the user's question verbatim, filler content |
5 — Exceptional: no reasonable improvement possible
4 — Good: minor nits only, no substantive gaps
3 — Adequate: meets the request but has a notable weakness on at least one axis
2 — Weak: has a clear gap that affects usability or correctness
1 — Poor: fundamentally misses the request or contains significant errorsEvery score below 5 MUST cite specific evidence. A score of 3 cannot just say "could be better" — it must say exactly what is missing or wrong. The mantra: "Show the gap, don't just name it."
Gather what you'll evaluate:
- The original user request (read back from conversation)
- Your final response/output (the deliverable)
- Any tool outputs that verify correctness (test results, exit codes, lint output)
- Any user feedback received during the task (corrections, "try again", "that's not right")Work through the 5 axes one at a time. For each:
Do NOT average the scores in your head first and then work backwards. Score each axis fresh.
Use the template from templates/evaluation-report.md. The report must include:
- One-line summary
- 5-axis scorecard (score + evidence per axis)
- Overall score (simple average, rounded to 1 decimal)
- 1-3 specific improvements ranked by impact
- Self-check: "Would the user agree with this assessment?"If any axis scored 3 or below:
Task: Add retry logic to HTTP client
Scorecard:
Accuracy: 5 — All API calls correct. Verified: retries use
exponential backoff. No hallucinated methods.
Completeness: 4 — Covered happy path + 3 error cases. Missing:
timeout handling for hung connections.
Clarity: 5 — Code comments explain backoff formula.
PR description links to incident that motivated this.
Actionability:5 — Single merge. No follow-up tasks. Tests pass.
Conciseness: 4 — 47 lines total. The retry loop could be extracted
into a helper to drop ~8 lines.
Overall: 4.6 — One gap (timeout handling). Fix before merging.Task: Add retry logic to HTTP client
Scorecard:
Accuracy: 2 — Used urllib3 which doesn't match our
httpx-based codebase. Wrong library.
Completeness: 3 — Works for GET. POST/PUT not handled (user
said "all HTTP requests").
Clarity: 4 — Code is readable. Good variable names.
Actionability:2 — "Add tests" mentioned but no test file created.
User has to write tests before merging.
Conciseness: 3 — 120 lines. The retry config is duplicated in
3 places instead of one shared RetryConfig object.
Overall: 2.8 — Wrong library used. Needs httpx rewrite.
Fix accuracy first (switch to httpx), then extend to all
HTTP methods, then consolidate config.FAIL: Accuracy: 5 — All good.
Completeness: 5 — Everything covered.
Clarity: 5 — Clear.No evidence cited. This is self-congratulation, not evaluation. A real 5 requires proving there's nothing to improve.
FAIL: Completeness: 2 — Didn't handle WebSocket connections or
gRPC streaming (user didn't ask for these)Only evaluate against what the user actually requested, not what you could have additionally built.
FAIL: "As I said earlier, this approach is wrong. Score: 1"The evaluation is about the delivered output, not about re-arguing design decisions that were already made. If the approach was wrong, that should have been caught before delivery.
FAIL: "Score: 3. I don't like Python decorators.""Don't like" is not evidence. Cite a concrete readability, testability, or correctness concern, or leave the score at 4+.
agent-eval — Head-to-head comparison of different coding agents on benchmark tasksverification-loop — Systematic verification of outputs against expected resultssecurity-review — Security-focused code review checklist© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in skills/agent-self-evaluation of affaan-m/ECC.
Open the folder on GitHubat commit ef648e0
Agent Self Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Self Evaluation this skillaffaan-m/ECC | 276k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Web Interface Guidelines Reviewervercel-labs/openreview | 1.7k | 97 repos | ~308 | Automated safety check: Pass | None | |
| Accessibility Reviewmarkmead/hyperui | 12k | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Web Animation DesignbaptisteArno/typebot.io | 11k | 2 repos | ~2.7k | Automated safety check: Pass | Custom licence | |
| Accessibility Fixeribelick/ui-skills | 9.5k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Wcag Audit PatternsvmDeshpande/ai-agent-automation | 178 | 11 repos | ~610 | Automated safety check: Pass | Apache-2.0 |
vercel-labs/openreview
Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my…
markmead/hyperui
Run a WCAG 2.1 AA accessibility audit on a design or page. An agent skill from markmead/hyperui.
baptisteArno/typebot.io
Guides easing, timing and animation choices for UI motion, based on a web animation course, and reviews existing animations in a before-and-after table.
ibelick/ui-skills
Audits and fixes HTML accessibility problems such as ARIA labels, keyboard navigation, focus management, contrast and form errors with minimal changes.
vmDeshpande/ai-agent-automation
Conduct WCAG 2.2 accessibility audits with automated testing, manual verification, and remediation guidance.
ibelick/ui-skills
Applies a fixed set of UI rules for stack, components, interaction, animation, typography and layout, or reviews a file against them with concrete fixes.
affaan-m/ECC
Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.
affaan-m/ECC
Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.
affaan-m/ECC
Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.
affaan-m/ECC
Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.
affaan-m/ECC
Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.
affaan-m/ECC
Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.
Categories
Use after completing any non-trivial task. An agent skill from affaan-m/ECC. Agent Self Evaluation is an agent skill from affaan-m/ECC. Use after completing any non-trivial task.
Agent Self Evaluation fits situations like: tasks that involve Accessibility.
Run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a claude-code`. Or copy the skill folder (skills/agent-self-evaluation in affaan-m/ECC) into .claude/skills/agent-self-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a codex`. Or copy the skill folder (skills/agent-self-evaluation in affaan-m/ECC) into .agents/skills/agent-self-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-self-evaluation, .gemini/skills/agent-self-evaluation, .github/skills/agent-self-evaluation and .opencode/skills/agent-self-evaluation in your project.
Going by SKILL.md and its folder, Agent Self Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Agent Self Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agent Self Evaluation: Web Interface Guidelines Reviewer (vercel-labs/openreview, 1.7k stars), Accessibility Review (markmead/hyperui, 12k stars), Web Animation Design (baptisteArno/typebot.io, 11k stars) and Accessibility Fixer (ibelick/ui-skills, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.
Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.