Agent skill

Agent Eval

by affaan-m in affaan-m/ECC

编码代理(Claude Code、Aider、Codex等)在自定义任务上的直接比较,包含通过率、成本、时间和一致性指标

MITAuto-check passedDevelopment

Install Agent Eval

skills CLI
$ npx skills add affaan-m/ECC --skill agent-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC agent-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/zh-CN/skills/agent-eval .claude/skills/agent-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-eval
GitHub stars
276k
Used in
1 other repo
Token cost
~730 tokens
SKILL.md length
88 words
Files
1
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

编码代理(Claude Code、Aider、Codex等)在自定义任务上的直接比较,包含通过率、成本、时间和一致性指标

  • Works in 3 steps: 定义任务 → 运行代理 → 比较结果
  • Development work in your project
  • SKILL.md covers 何时使用, 安装, 核心概念 and 工作流程, plus 3 more sections
  • Calls pip; reaches github.com

What it does

Agent Eval is an agent skill from affaan-m/ECC. 编码代理(Claude Code、Aider、Codex等)在自定义任务上的直接比较,包含通过率、成本、时间和一致性指标

Its SKILL.md is about 730 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. It works with Git. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Development work in your project

Example prompts

  • “/agent-eval”

Requirements

  • Python 3
  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 定义任务
  2. 运行代理
  3. 比较结果

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Eval loads about 730 tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 88 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~730

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 88 words, ~730 tokens.

Download SKILL.mdSave it as .claude/skills/agent-eval/SKILL.md (or your agent's skills folder).
name
agent-eval
description
编码代理(Claude Code、Aider、Codex等)在自定义任务上的直接比较,包含通过率、成本、时间和一致性指标
origin
ECC
tools
Read, Write, Edit, Bash, Grep, Glob

Agent Eval 技能

一个轻量级 CLI 工具,用于在可复现的任务上对编码代理进行头对头比较。每个“哪个编码代理最好?”的比较都基于感觉——本工具将其系统化。

何时使用

  • 在你自己的代码库上比较编码代理(Claude Code、Aider、Codex 等)
  • 在采用新工具或模型之前衡量代理性能
  • 当代理更新其模型或工具时运行回归检查
  • 为团队做出数据支持的代理选择决策

安装

bash
# pinned to v0.1.0 — latest stable commit
pip install git+https://github.com/joaquinhuigomez/agent-eval.git@6d062a2f5cda6ea443bf5d458d361892c04e749b

核心概念

YAML 任务定义

以声明方式定义任务。每个任务指定要做什么、要修改哪些文件以及如何判断成功:

yaml
name: add-retry-logic
description: Add exponential backoff retry to the HTTP client
repo: ./my-project
files:
  - src/http_client.py
prompt: |
  Add retry logic with exponential backoff to all HTTP requests.
  Max 3 retries. Initial delay 1s, max delay 30s.
judge:
  - type: pytest
    command: pytest tests/test_http_client.py -v
  - type: grep
    pattern: "exponential_backoff|retry"
    files: src/http_client.py
commit: "abc1234"  # pin to specific commit for reproducibility
Git 工作树隔离

每个代理运行都获得自己的 git 工作树——无需 Docker。这提供了可复现的隔离,使得代理之间不会相互干扰或损坏基础仓库。

收集的指标
指标衡量内容
通过率代理生成的代码是否通过了判断?
成本每个任务的 API 花费(如果可用)
时间完成所需的挂钟秒数
一致性跨重复运行的通过率(例如,3/3 = 100%)

工作流程

1. 定义任务

创建一个 tasks/ 目录,其中包含 YAML 文件,每个任务一个文件:

bash
mkdir tasks
# Write task definitions (see template above)
2. 运行代理

针对你的任务执行代理:

bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3

每次运行:

  1. 从指定的提交创建一个新的 git 工作树
  2. 将提示交给代理
  3. 运行判断标准
  4. 记录通过/失败、成本和时间
3. 比较结果

生成比较报告:

bash
agent-eval report --format table
Task: add-retry-logic (3 runs each)
┌──────────────┬───────────┬────────┬────────┬─────────────┐
│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │
├──────────────┼───────────┼────────┼────────┼─────────────┤
│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        │
│ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │
└──────────────┴───────────┴────────┴────────┴─────────────┘

判断类型

基于代码(确定性)
yaml
judge:
  - type: pytest
    command: pytest tests/ -v
  - type: command
    command: npm run build
基于模式
yaml
judge:
  - type: grep
    pattern: "class.*Retry"
    files: src/**/*.py
基于模型(LLM 作为判断器)
yaml
judge:
  - type: llm
    prompt: |
      Does this implementation correctly handle exponential backoff?
      Check for: max retries, increasing delays, jitter.

最佳实践

  • 从 3-5 个任务开始,这些任务代表你的真实工作负载,而非玩具示例
  • 每个代理至少运行 3 次试验以捕捉方差——代理是非确定性的
  • 在你的任务 YAML 中固定提交,以便结果在数天/数周内可复现
  • 每个任务至少包含一个确定性判断器(测试、构建)——LLM 判断器会增加噪音
  • 跟踪成本与通过率——一个通过率 95% 但成本高出 10 倍的代理可能不是正确的选择
  • 对你的任务定义进行版本控制——它们是测试夹具,应将其视为代码

链接

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/zh-CN/skills/agent-eval of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Eval this skillaffaan-m/ECC276k1 repos~730Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
Contributor-First PR MergeHKUDS/OpenHarness16k1 repos~847Automated safety check: PassMIT
Finishing A Development Branchfarm-fe/farm5.6k34 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Merges external GitHub pull requests while keeping the original author credited, and fixes conflicts after the merge instead of rewriting the contribution.

    16k GitHub starsUsed in 1 repo~847 tokens
    DevelopmentAuto-check passed
  • A skill your agent uses when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for…

    5.6k GitHub starsUsed in 34 repos~1.8k tokens
    DevelopmentAuto-check passed
  • Moves a package from another TryGhost repository into Ghost as an internal workspace package while keeping its Git history, with checkpoints for the steps that need an administrator.

    56k GitHub stars~3.8k tokensUpdated today
    DevelopmentAuto-check passed

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Works with

Categories

Questions about Agent Eval

What does Agent Eval do?

编码代理(Claude Code、Aider、Codex等)在自定义任务上的直接比较,包含通过率、成本、时间和一致性指标. Agent Eval is an agent skill from affaan-m/ECC.

When should I use Agent Eval?

Agent Eval fits situations like: development work in your project.

How do I install Agent Eval in Claude Code?

Run `npx skills add affaan-m/ECC --skill agent-eval -a claude-code`. Or copy the skill folder (docs/zh-CN/skills/agent-eval in affaan-m/ECC) into .claude/skills/agent-eval in your project. Claude Code loads it when a task matches its description.

How do I install Agent Eval in Codex?

Run `npx skills add affaan-m/ECC --skill agent-eval -a codex`. Or copy the skill folder (docs/zh-CN/skills/agent-eval in affaan-m/ECC) into .agents/skills/agent-eval in your project. Codex loads it when a task matches its description.

Can I use Agent Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill agent-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-eval, .gemini/skills/agent-eval, .github/skills/agent-eval and .opencode/skills/agent-eval in your project.

What does Agent Eval need to run?

Going by SKILL.md and its folder, Agent Eval needs the command-line tools its instructions call (pip). Our summary lists: Python 3; Docker.

Does Agent Eval access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Agent Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Eval use?

Agent Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Eval use?

About 730 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Eval?

Skills that share tags, products or a category with Agent Eval: Finishing a Development Branch (obra/superpowers, 297k stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Code Design Rationale Investigator (cursor/plugins, 10k stars) and Contributor-First PR Merge (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Eval?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.