Agent skill

Agent Self Evaluation

by affaan-m in affaan-m/ECC

Use after completing any non-trivial task. An agent skill from affaan-m/ECC.

MITAuto-check passedFrontend & Design

Install Agent Self Evaluation

skills CLI
$ npx skills add affaan-m/ECC --skill agent-self-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC agent-self-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-self-evaluation .claude/skills/agent-self-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-self-evaluation
GitHub stars
276k
Token cost
~1.9k tokens
SKILL.md length
665 words
Files
7 (incl. scripts, references)
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

Use after completing any non-trivial task. An agent skill from affaan-m/ECC.

  • Works in 4 steps: Collect the Raw Material → Score Each Axis Independently → Produce the Evaluation Report → …
  • Tasks that involve Accessibility
  • SKILL.md covers When to Activate, Core Concepts, Workflow and Code Examples, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Agent Self Evaluation is an agent skill from affaan-m/ECC. Use after completing any non-trivial task. The agent self-rates its output on 5 axes — accuracy, completeness, clarity, actionability, conciseness — with concrete evidence per criterion. Produces a structured 1-5 scorecard with specific improvement suggestions.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `examples/high-score-example.md`, `examples/low-score-example.md` and `references/evaluation-criteria.md`).

It sits in Frontend & Design, covering Accessibility. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Tasks that involve Accessibility

Example prompts

  • “/agent-self-evaluation”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Collect the Raw Material
  2. Score Each Axis Independently
  3. Produce the Evaluation Report
  4. Apply the Improvement

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Self Evaluation loads about 1.9k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 665 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 665 words, ~1,890 tokens.

Download SKILL.mdSave it as .claude/skills/agent-self-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
agent-self-evaluation
description
Use after completing any non-trivial task. The agent self-rates its output on 5 axes — accuracy, completeness, clarity, actionability, conciseness — with concrete evidence per criterion. Produces a structured 1-5 scorecard with specific improvement suggestions.
origin
ECC

Agent Self-Evaluation

After completing a complex task, the agent pauses to rate its own output against a structured 5-axis rubric. This is NOT a pass/fail gate — it's a deliberate reflection step that catches omissions, flags overconfidence, and surface areas for improvement before the user has to.

When to Activate

  • After writing code that spans 3+ files or 50+ lines
  • After completing a multi-step workflow (implement → test → review)
  • After a debugging session that involved 3+ attempts
  • After producing a design document, architecture decision, or written analysis
  • When the user asks "how good was that?" or "rate yourself"
  • At the end of any session Stop hook (if configured — see references/hook-integration.md)

Core Concepts

The 5 Evaluation Axes
AxisQuestionWhat it catches
AccuracyAre the facts, claims, and outputs correct?Hallucinations, wrong API names, incorrect syntax, false statements
CompletenessDid it cover everything the user asked for?Missed edge cases, unhandled error paths, forgotten requirements, skipped subtasks
ClarityIs the explanation understandable and well-structured?Confusing explanations, jargon without definition, missing context, rambling
ActionabilityCan the user act on the output immediately?Vague suggestions, missing steps, "you should X" without showing how, no verification path
ConcisenessDid it use the minimum words/tokens needed?Redundancy, over-explanation, repeating the user's question verbatim, filler content
Scoring Scale
5 — Exceptional: no reasonable improvement possible
4 — Good: minor nits only, no substantive gaps
3 — Adequate: meets the request but has a notable weakness on at least one axis
2 — Weak: has a clear gap that affects usability or correctness
1 — Poor: fundamentally misses the request or contains significant errors
The Evidence Rule

Every score below 5 MUST cite specific evidence. A score of 3 cannot just say "could be better" — it must say exactly what is missing or wrong. The mantra: "Show the gap, don't just name it."

Workflow

Step 1: Collect the Raw Material

Gather what you'll evaluate:

- The original user request (read back from conversation)
- Your final response/output (the deliverable)
- Any tool outputs that verify correctness (test results, exit codes, lint output)
- Any user feedback received during the task (corrections, "try again", "that's not right")
Step 2: Score Each Axis Independently

Work through the 5 axes one at a time. For each:

  1. Read the axis question
  2. Find evidence (or lack of evidence) in the output
  3. Assign a score 1-5
  4. If score < 5, write a one-sentence improvement note citing the gap

Do NOT average the scores in your head first and then work backwards. Score each axis fresh.

Step 3: Produce the Evaluation Report

Use the template from templates/evaluation-report.md. The report must include:

- One-line summary
- 5-axis scorecard (score + evidence per axis)
- Overall score (simple average, rounded to 1 decimal)
- 1-3 specific improvements ranked by impact
- Self-check: "Would the user agree with this assessment?"
Step 4: Apply the Improvement

If any axis scored 3 or below:

  1. State what you would do differently
  2. If the gap is fixable in < 30 seconds (missing link, unclear phrasing), fix it now
  3. If the gap requires rework, flag it explicitly: "This axis scored [reason] because [evidence]. Re-running with [specific fix] would likely raise it to [score]."
Show full SKILL.md (264 more words)Show less

Code Examples

Example: Good Evaluation (Score 4+)
Task: Add retry logic to HTTP client

Scorecard:
  Accuracy:    5 — All API calls correct. Verified: retries use
                  exponential backoff. No hallucinated methods.
  Completeness: 4 — Covered happy path + 3 error cases. Missing:
                  timeout handling for hung connections.
  Clarity:      5 — Code comments explain backoff formula.
                  PR description links to incident that motivated this.
  Actionability:5 — Single merge. No follow-up tasks. Tests pass.
  Conciseness:  4 — 47 lines total. The retry loop could be extracted
                  into a helper to drop ~8 lines.

Overall: 4.6 — One gap (timeout handling). Fix before merging.
Example: Weak Evaluation (Score 2-3)
Task: Add retry logic to HTTP client

Scorecard:
  Accuracy:    2 — Used urllib3 which doesn't match our
                  httpx-based codebase. Wrong library.
  Completeness: 3 — Works for GET. POST/PUT not handled (user
                  said "all HTTP requests").
  Clarity:      4 — Code is readable. Good variable names.
  Actionability:2 — "Add tests" mentioned but no test file created.
                  User has to write tests before merging.
  Conciseness:  3 — 120 lines. The retry config is duplicated in
                  3 places instead of one shared RetryConfig object.

Overall: 2.8 — Wrong library used. Needs httpx rewrite.
  Fix accuracy first (switch to httpx), then extend to all
  HTTP methods, then consolidate config.

Anti-Patterns

"Everything is a 5"
FAIL: Accuracy:    5 — All good.
   Completeness: 5 — Everything covered.
   Clarity:      5 — Clear.

No evidence cited. This is self-congratulation, not evaluation. A real 5 requires proving there's nothing to improve.

Over-penalizing for scope creep
FAIL: Completeness: 2 — Didn't handle WebSocket connections or
   gRPC streaming (user didn't ask for these)

Only evaluate against what the user actually requested, not what you could have additionally built.

Using the evaluation to re-litigate
FAIL: "As I said earlier, this approach is wrong. Score: 1"

The evaluation is about the delivered output, not about re-arguing design decisions that were already made. If the approach was wrong, that should have been caught before delivery.

Mixing personal preference with objective gaps
FAIL: "Score: 3. I don't like Python decorators."

"Don't like" is not evidence. Cite a concrete readability, testability, or correctness concern, or leave the score at 4+.

Best Practices

  • Evaluate the output, not the process. The user cares about what you delivered, not how many iterations you took.
  • One improvement per weak axis. Don't list 5 things for one axis — pick the highest-impact gap.
  • Tie improvements to user impact. "Missing error handling means the user's API call will crash silently" beats "add error handling."
  • Be specific about what 'fixed' looks like. "Re-run with httpx transport configured for retries" beats "fix the library issue."
  • Use tool outputs as evidence. If tests passed, cite them. If lint is clean, cite it. Don't guess — grep for the proof.
  • If you can't find any gaps, try harder. A perfect score across all 5 axes is rare. Ask: "If I were the user, what would annoy me about this output?"
  • agent-eval — Head-to-head comparison of different coding agents on benchmark tasks
  • verification-loop — Systematic verification of outputs against expected results
  • security-review — Security-focused code review checklist

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/agent-self-evaluation of affaan-m/ECC.

  • SKILL.md
  • examples/high-score-example.md
  • examples/low-score-example.md
  • references/evaluation-criteria.md
  • references/hook-integration.md
  • scripts/evaluate.py
  • templates/evaluation-report.md

Open the folder on GitHubat commit ef648e0

Compare with similar skills

Agent Self Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Self Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Self Evaluation this skillaffaan-m/ECC276k—~1.9kAutomated safety check: PassMIT
Web Interface Guidelines Reviewervercel-labs/openreview1.7k97 repos~308Automated safety check: PassNone
Accessibility Reviewmarkmead/hyperui12k1 repos~1.1kAutomated safety check: PassMIT
Web Animation DesignbaptisteArno/typebot.io11k2 repos~2.7kAutomated safety check: PassCustom licence
Accessibility Fixeribelick/ui-skills9.5k4 repos~1.2kAutomated safety check: PassMIT
Wcag Audit PatternsvmDeshpande/ai-agent-automation17811 repos~610Automated safety check: PassApache-2.0

Similar skills

  • Web Interface Guidelines Reviewer

    vercel-labs/openreview

    Official

    Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my…

    1.7k GitHub starsUsed in 97 repos~308 tokens
    Frontend & DesignAuto-check passed
  • Accessibility Review

    markmead/hyperui

    Run a WCAG 2.1 AA accessibility audit on a design or page. An agent skill from markmead/hyperui.

    12k GitHub starsUsed in 1 repo~1.1k tokens
    Frontend & DesignAuto-check passed
  • Web Animation Design

    baptisteArno/typebot.io

    Guides easing, timing and animation choices for UI motion, based on a web animation course, and reviews existing animations in a before-and-after table.

    11k GitHub starsUsed in 2 repos~2.7k tokens
    Frontend & DesignAuto-check passed
  • Accessibility Fixer

    ibelick/ui-skills

    Audits and fixes HTML accessibility problems such as ARIA labels, keyboard navigation, focus management, contrast and form errors with minimal changes.

    9.5k GitHub starsUsed in 4 repos~1.2k tokens
    Frontend & DesignAuto-check passed
  • Wcag Audit Patterns

    vmDeshpande/ai-agent-automation

    Conduct WCAG 2.2 accessibility audits with automated testing, manual verification, and remediation guidance.

    178 GitHub starsUsed in 11 repos~610 tokens
    Frontend & DesignAuto-check passed
  • Baseline UI

    ibelick/ui-skills

    Applies a fixed set of UI rules for stack, components, interaction, animation, typography and layout, or reviews a file against them with concrete fixes.

    9.5k GitHub starsUsed in 8 repos~855 tokens
    Frontend & DesignAuto-check passed

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Questions about Agent Self Evaluation

What does Agent Self Evaluation do?

Use after completing any non-trivial task. An agent skill from affaan-m/ECC. Agent Self Evaluation is an agent skill from affaan-m/ECC. Use after completing any non-trivial task.

When should I use Agent Self Evaluation?

Agent Self Evaluation fits situations like: tasks that involve Accessibility.

How do I install Agent Self Evaluation in Claude Code?

Run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a claude-code`. Or copy the skill folder (skills/agent-self-evaluation in affaan-m/ECC) into .claude/skills/agent-self-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Agent Self Evaluation in Codex?

Run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a codex`. Or copy the skill folder (skills/agent-self-evaluation in affaan-m/ECC) into .agents/skills/agent-self-evaluation in your project. Codex loads it when a task matches its description.

Can I use Agent Self Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill agent-self-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-self-evaluation, .gemini/skills/agent-self-evaluation, .github/skills/agent-self-evaluation and .opencode/skills/agent-self-evaluation in your project.

What does Agent Self Evaluation need to run?

Going by SKILL.md and its folder, Agent Self Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Agent Self Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Self Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Self Evaluation use?

Agent Self Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Self Evaluation use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Agent Self Evaluation?

Skills that share tags, products or a category with Agent Self Evaluation: Web Interface Guidelines Reviewer (vercel-labs/openreview, 1.7k stars), Accessibility Review (markmead/hyperui, 12k stars), Web Animation Design (baptisteArno/typebot.io, 11k stars) and Accessibility Fixer (ibelick/ui-skills, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Self Evaluation?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.