Agent skill

Agent Eval

by affaan-m in affaan-m/ECC

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics.

MITAuto-check passedAI & LLM Engineering

Install Agent Eval

skills CLI
$ npx skills add affaan-m/ECC --skill agent-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC agent-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-eval .claude/skills/agent-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-eval
GitHub stars
276k
Used in
3 other repos
Token cost
~1.1k tokens
SKILL.md length
328 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics.

  • Works in 3 steps: Define Tasks → Run Agents → Compare Results
  • Choosing between coding agents
  • SKILL.md covers When to Activate, Installation, Core Concepts and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Eval is an agent skill from affaan-m/ECC. Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with Git. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Choosing between coding agents
  • A change to an agent setup needs measured pass rate
  • Time rather than an impression

Example prompts

  • “/agent-eval”

Requirements

  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Define Tasks
  2. Run Agents
  3. Compare Results

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Eval loads about 1.1k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 328 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 328 words, ~1,072 tokens.

Download SKILL.mdSave it as .claude/skills/agent-eval/SKILL.md (or your agent's skills folder).
name
agent-eval
description
Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.
license
MIT
metadata.origin
ECC
tools
Read, Write, Edit, Bash, Grep, Glob

Agent Eval Skill

A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.

When to Activate

  • Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
  • Measuring agent performance before adopting a new tool or model
  • Running regression checks when an agent updates its model or tooling
  • Producing data-backed agent selection decisions for a team

Installation

Note: Install agent-eval from its repository after reviewing the source.

Core Concepts

YAML Task Definitions

Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:

yaml
name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
  - src/http_client.py
prompt: |
  Add retry logic with exponential backoff to all HTTP requests.
  Max 3 retries. Initial delay 1s, max delay 30s.
judge:
  - type: pytest
    command: pytest tests/test_http_client.py -v
  - type: grep
    pattern: "exponential_backoff|retry"
    files: src/http_client.py
commit: "abc1234"  # pin to specific commit for reproducibility
Git Worktree Isolation

Each agent run gets its own git worktree — no Docker required. This provides reproducibility isolation so agents cannot interfere with each other or corrupt the base repo.

Metrics Collected
MetricWhat It Measures
Pass rateDid the agent produce code that passes the judge?
CostAPI spend per task (when available)
TimeWall-clock seconds to completion
ConsistencyPass rate across repeated runs (e.g., 3/3 = 100%)

Workflow

1. Define Tasks

Create a tasks/ directory with YAML files, one per task:

bash
mkdir tasks
# Write task definitions (see template above)
2. Run Agents

Execute agents against your tasks:

bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3

Each run:

  1. Creates a fresh git worktree from the specified commit
  2. Hands the prompt to the agent
  3. Runs the judge criteria
  4. Records pass/fail, cost, and time
3. Compare Results

Generate a comparison report:

bash
agent-eval report --format table
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        │
│ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │
└──────────────┴───────────┴────────┴────────┴─────────────┘

Judge Types

Code-Based (deterministic)
yaml
judge:
  - type: pytest
    command: pytest tests/ -v
  - type: command
    command: npm run build
Pattern-Based
yaml
judge:
  - type: grep
    pattern: "class.*Retry"
    files: src/**/*.py
Model-Based (LLM-as-judge)
yaml
judge:
  - type: llm
    prompt: |
      Does this implementation correctly handle exponential backoff?
      Check for: max retries, increasing delays, jitter.

Best Practices

  • Start with 3-5 tasks that represent your real workload, not toy examples
  • Run at least 3 trials per agent to capture variance — agents are non-deterministic
  • Pin the commit in your task YAML so results are reproducible across days/weeks
  • Include at least one deterministic judge (tests, build) per task — LLM judges add noise
  • Track cost alongside pass rate — a 95% agent at 10x the cost may not be the right choice
  • Version your task definitions — they are test fixtures, treat them as code

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-eval of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Used in 3 other repositories

We found 9 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Agent Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Eval this skillaffaan-m/ECC276k3 repos~1.1kAutomated safety check: PassMIT
Opt Trace Miningalibaba/atrex-kernel-agent167—~4.4kAutomated safety check: PassApache-2.0
Ownmem Dashboardgrpcer/ownmem423—~563Automated safety check: PassApache-2.0
Use Shared Credentialyc-software/qm15k—~1.3kAutomated safety check: PassMIT
Cc Update ReviewChachamaru127/claude-code-harness3.2k—~1.6kAutomated safety check: NotesMIT
Gltf Asset Optimizationelodin-sys/elodin547—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Opt Trace Mining

    alibaba/atrex-kernel-agent

    Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…

    167 GitHub stars~4.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ownmem Dashboard

    grpcer/ownmem

    Open OwnMem Console, the local dashboard for this repository's memory.

    423 GitHub stars~563 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Use Shared Credential

    yc-software/qm

    When you need to call a service the org has a SHARED credential for (a SERP/search key, a paid-API key, a data-vendor feed, a Git remote) — and the platform has told you that credential is available…

    15k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cc Update Review

    Chachamaru127/claude-code-harness

    Quality guardrail for Claude/Codex update integration. An agent skill from Chachamaru127/claude-code-harness.

    3.2k GitHub stars~1.6k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Gltf Asset Optimization

    elodin-sys/elodin

    Reduce the size of glTF/GLB 3D assets to cut Git LFS bandwidth/storage while keeping them loadable by the editor's Bevy 0.18 glTF loader.

    547 GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Reliability Concurrency

    ChatbotXIO/ChatbotX

    A skill your agent uses when writing code that runs concurrently in ChatbotX — BullMQ worker consumers, sharded DB migrations, embedding replace-writes, or any multi-step operation that could be…

    885 GitHub stars~735 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 683 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    276k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed

Works with

Questions about Agent Eval

What does Agent Eval do?

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Agent Eval is an agent skill from affaan-m/ECC.) on custom tasks with pass rate, cost, time, and consistency metrics.

When should I use Agent Eval?

Agent Eval fits situations like: choosing between coding agents; A change to an agent setup needs measured pass rate; time rather than an impression.

How do I install Agent Eval in Claude Code?

Run `npx skills add affaan-m/ECC --skill agent-eval -a claude-code`. Or copy the skill folder (skills/agent-eval in affaan-m/ECC) into .claude/skills/agent-eval in your project. Claude Code loads it when a task matches its description.

How do I install Agent Eval in Codex?

Run `npx skills add affaan-m/ECC --skill agent-eval -a codex`. Or copy the skill folder (skills/agent-eval in affaan-m/ECC) into .agents/skills/agent-eval in your project. Codex loads it when a task matches its description.

Can I use Agent Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill agent-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-eval, .gemini/skills/agent-eval, .github/skills/agent-eval and .opencode/skills/agent-eval in your project.

What does Agent Eval need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Eval is instructions for the agent only. Our summary lists: Docker.

Does Agent Eval access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Agent Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Eval use?

Agent Eval is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Eval use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Eval?

Skills that share tags, products or a category with Agent Eval: Opt Trace Mining (alibaba/atrex-kernel-agent, 167 stars), Ownmem Dashboard (grpcer/ownmem, 423 stars), Use Shared Credential (yc-software/qm, 15k stars) and Cc Update Review (Chachamaru127/claude-code-harness, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Eval?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.