Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

MITAuto-check passedAI & LLM Engineering

Install Eval

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/agenthub/skills/eval .claude/skills/eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval
GitHub stars
28k
Used in
1 other repo
Token cost
~618 tokens
SKILL.md length
169 words
Files
1
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

  • Works in 3 steps: Get the diff: git diff… → Read the agent's result post from… → Compare all diffs and rank by
  • The user runs /hub:eval
  • SKILL.md covers Usage, What It Does and After Eval
  • Calls python and git

What it does

Eval is an agent skill from alirezarezvani/claude-skills. Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.

Its SKILL.md is about 620 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • The user runs /hub:eval
  • Pick a winner among completed AgentHub agents

Example prompts

  • “/eval”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Get the diff: git diff {base_branch}...{agent_branch}
  2. Read the agent's result post from .agenthub/board/results/agent-{i}-result.md
  3. Compare all diffs and rank by

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval loads about 618 tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 169 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~618

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 169 words, ~618 tokens.

Download SKILL.mdSave it as .claude/skills/eval/SKILL.md (or your agent's skills folder).
name
eval
description
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
command
/hub:eval

/hub:eval — Evaluate Agent Results

Rank all agent results for a session. Supports metric-based evaluation (run a command), LLM judge (compare diffs), or hybrid.

Usage

/hub:eval                           # Eval latest session using configured criteria
/hub:eval 20260317-143022           # Eval specific session
/hub:eval --judge                   # Force LLM judge mode (ignore metric config)

What It Does

Metric Mode (eval command configured)

Run the evaluation command in each agent's worktree:

bash
python {skill_path}/scripts/result_ranker.py \
  --session {session-id} \
  --eval-cmd "{eval_cmd}" \
  --metric {metric} --direction {direction}

Output:

RANK  AGENT       METRIC      DELTA      FILES
1     agent-2     142ms       -38ms      2
2     agent-1     165ms       -15ms      3
3     agent-3     190ms       +10ms      1

Winner: agent-2 (142ms)
LLM Judge Mode (no eval command, or --judge flag)

For each agent:

  1. Get the diff: git diff {base_branch}...{agent_branch}
  2. Read the agent's result post from .agenthub/board/results/agent-{i}-result.md
  3. Compare all diffs and rank by:
    • Correctness — Does it solve the task?
    • Simplicity — Fewer lines changed is better (when equal correctness)
    • Quality — Clean execution, good structure, no regressions

Present rankings with justification.

Example LLM judge output for a content task:

RANK  AGENT    VERDICT                               WORD COUNT
1     agent-1  Strong narrative, clear CTA            1480
2     agent-3  Good data points, weak intro           1520
3     agent-2  Generic tone, no differentiation       1350

Winner: agent-1 (strongest narrative arc and call-to-action)
Hybrid Mode
  1. Run metric evaluation first
  2. If top agents are within 10% of each other, use LLM judge to break ties
  3. Present both metric and qualitative rankings

After Eval

  1. Update session state:
bash
python {skill_path}/scripts/session_manager.py --update {session-id} --state evaluating
  1. Tell the user:
    • Ranked results with winner highlighted
    • Next step: /hub:merge to merge the winner
    • Or /hub:merge {session-id} --agent {winner} to be explicit

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in engineering/agenthub/skills/eval of alirezarezvani/claude-skills.

Open the folder on GitHubat commit 19392f7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in alirezarezvani/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval this skillalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about Eval

What does Eval do?

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Eval is an agent skill from alirezarezvani/claude-skills. Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

When should I use Eval?

Eval fits situations like: the user runs /hub:eval; pick a winner among completed AgentHub agents.

How do I install Eval in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill eval -a claude-code`. Or copy the skill folder (engineering/agenthub/skills/eval in alirezarezvani/claude-skills) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.

How do I install Eval in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill eval -a codex`. Or copy the skill folder (engineering/agenthub/skills/eval in alirezarezvani/claude-skills) into .agents/skills/eval in your project. Codex loads it when a task matches its description.

Can I use Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.

What does Eval need to run?

Going by SKILL.md and its folder, Eval needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Eval access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval use?

Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval use?

About 618 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval?

Skills that share tags, products or a category with Eval: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,788 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.