Agent skill

Agentic Evaluation Framework

by borghei in borghei/Claude-Skills

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

MITAuto-check passedEducation

Install Agentic Evaluation Framework

skills CLI
$ npx skills add borghei/Claude-Skills --skill agentic-evaluation-framework -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills agentic-evaluation-framework --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/agentic-evaluation-framework .claude/skills/agentic-evaluation-framework && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-evaluation-framework
GitHub stars
874
Token cost
~1.9k tokens
SKILL.md length
960 words
Files
5 (incl. scripts, references)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

  • Works in 2 steps: Build and calibrate an absolute-scoring… → Rank model/prompt variants by pairwise…
  • Asks to evaluate LLM output quality
  • SKILL.md covers Overview, Clarify First, Quick Start and Tools Overview, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Agentic Evaluation Framework is an agent skill from borghei/Claude-Skills. This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/eval-pitfalls.md`, `references/llm-judge-methodology.md` and `scripts/pairwise_ranking.py`).

It sits in Education, covering Quizzes and assessments and LLM evaluation. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Asks to evaluate LLM output quality
  • Set up LLM-as-judge
  • Build an eval rubric
  • Compare model outputs pairwise

Example prompts

  • “evaluate LLM output quality”
  • “set up LLM-as-judge”
  • “build an eval rubric”
  • “/agentic-evaluation-framework”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Build and calibrate an absolute-scoring rubric
  2. Rank model/prompt variants by pairwise comparison

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic Evaluation Framework loads about 1.9k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 54 tokens; SKILL.md has 960 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 960 words, ~1,887 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-evaluation-framework/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agentic-evaluation-framework
description
This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
ai-engineering
metadata.updated
2026-06-29
metadata.tags
evaluation, llm-as-judge, rubrics, agents, quality

Agentic Evaluation Framework

Category: Engineering Domain: AI Engineering

Overview

Design and run trustworthy evaluations for LLM and agent outputs: pick the right grading method (programmatic check, LLM-as-judge, or human review), write a scoring rubric that judges can apply consistently, rank competing variants by pairwise comparison, and watch for the biases that quietly corrupt judge scores — position bias, verbosity bias, and self-preference. The goal is an eval that you can trust enough to ship on: calibrated against human labels, cheap enough to run on every change, and tracked alongside cost and latency so you never trade quality away by accident. This skill is model- and vendor-agnostic: it reasons about the evaluation method, not any one provider's API, and its scripts aggregate scores you have already collected — they never call a model.

Clarify First

Before designing or running an evaluation, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • What "good" means — the dimensions you care about (accuracy, helpfulness, safety, format, tool-use) and their relative weight (defines the rubric criteria and weights)
  • Grading method — can a deterministic check decide it, do you need an LLM judge, or must a human review it? (selects programmatic vs rubric_scorer.py absolute scoring vs pairwise_ranking.py comparison vs human-in-the-loop)
  • Ground truth & budget — do you have human-labeled examples to calibrate the judge against, and what cost/latency per eval run is acceptable? (sets calibration plan and the quality/cost/latency budget)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions.

Quick Start

bash
cd engineering/agentic-evaluation-framework

# 1. Score outputs against a weighted rubric + check inter-rater agreement
python scripts/rubric_scorer.py --data rubric_scores.json

# 2. Rank competing variants from pairwise (A-vs-B) judgements
python scripts/pairwise_ranking.py --data pairwise_matches.json

# JSON output for piping into a dashboard or CI gate
python scripts/rubric_scorer.py --data rubric_scores.json --json

Tools Overview

ToolPurposeKey Flags
scripts/rubric_scorer.pyAggregate per-criterion scores into weighted totals, per-criterion means, pass/fail vs thresholds, and an inter-rater agreement metric--data, --json
scripts/pairwise_ranking.pyTurn head-to-head win/loss records into a ranking via Elo + Bradley-Terry, plus a win-rate matrix--data, --k, --base, --json

Both scripts: Python 3 standard library only, argparse CLI, --json and human-readable output. They compute over scores you provide and never call a model. Run --help for full usage.

Workflows

1. Build and calibrate an absolute-scoring rubric
  1. Translate "what good means" into 3-6 named criteria, each with a weight, a 1-5 (or 1-7) scale, and a written anchor for every scale point (see references/llm-judge-methodology.md).
  2. Have at least two graders — human, or human plus model — independently score a calibration set, and record the per-criterion scores as the rubric_scorer.py input JSON.
  3. Run rubric_scorer.py and read inter_rater_agreement: low agreement means the rubric is ambiguous, not that a grader is wrong — tighten the anchors and re-score before trusting any number.
  4. Once graders agree, treat the human scores as ground truth and check that the LLM judge's scores correlate; if not, revise the judge prompt or fall back to human review for that criterion.
  5. Wire the passing rubric into CI as a gate (--json → pass/fail), and re-run agreement periodically to catch judge drift.
2. Rank model/prompt variants by pairwise comparison
  1. When absolute scores are noisy, switch to pairwise: show the judge two outputs (A and B) for the same input and ask only "which is better?" — easier and more reliable than an absolute number.
  2. Mitigate position bias by running each pair in both orders (A,B and B,A) and counting a win only if it survives both; record outcomes as pairwise_ranking.py matches, using "winner": "tie" for disagreements.
  3. Run pairwise_ranking.py to get Elo and Bradley-Terry rankings plus the win-rate matrix; Bradley-Terry is order-independent and preferred for a fixed batch, Elo for a streaming sequence of matches.
  4. Inspect the win-rate matrix for intransitivity (A>B, B>C, but C>A) — a sign of an unreliable judge or genuinely tied variants; collect more matches or add human adjudication.
  5. Report the ranking next to cost and latency per variant so the "winner" is the best quality-per-dollar-per-second, not just the highest score.
Show full SKILL.md (330 more words)Show less

Reference Documentation

  • references/llm-judge-methodology.md — rubric design and scale anchoring; absolute vs pairwise scoring; the judge-bias catalog (position, verbosity, self-preference, sycophancy) with concrete mitigations; calibrating a judge against human labels; the eval feedback loop (collect → grade → analyze → fix → regression-gate); and the quality/cost/latency metrics to track together.
  • references/eval-pitfalls.md — the anti-patterns that make evals lie: single-grader rubrics, judging on the training set, gameable metrics, ignoring variance, optimizing the judge instead of the model, and the decision table for when to use a programmatic check vs an LLM judge vs human review.

Common Patterns

  • Cheapest valid grader wins — if a deterministic check (regex, JSON-schema, unit test, exact match) can decide it, use that; reach for an LLM judge only for fuzzy quality, and human review only for high-stakes or judge-calibration work.
  • Pairwise over absolute when scores are noisy — "which is better, A or B?" is more reliable than "rate this 1-5"; use absolute rubrics for thresholds/gates and pairwise for model selection.
  • Swap positions to kill position bias — always run each comparison in both orders and only count wins that survive both; a variant that only wins in position A is a judge artifact.
  • Length is not quality — strip or normalize for verbosity bias; a longer answer is not a better one, and judges systematically over-reward length unless you control for it.
  • Don't let a model grade its own homework — self-preference bias means a model favors its own outputs; use a different judge family from the model under test, or anchor on human labels.
  • Calibrate before you trust — a judge is only as good as its agreement with humans on a held-out set; measure that agreement first, then automate.
  • Track quality, cost, and latency as one number — a quality win that triples cost or latency may be a net loss; always report the three together so the tradeoff is explicit.
  • Evals are regression tests for prompts — freeze a labeled eval set, gate every prompt/model change on it, and grow the set from production failures you find.

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in engineering/agentic-evaluation-framework of borghei/Claude-Skills.

  • SKILL.md
  • references/eval-pitfalls.md
  • references/llm-judge-methodology.md
  • scripts/pairwise_ranking.py
  • scripts/rubric_scorer.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Agentic Evaluation Framework next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic Evaluation Framework compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic Evaluation Framework this skillborghei/Claude-Skills874—~1.9kAutomated safety check: PassMIT
Woo AI Smokewoocommerce/woocommerce-ios3581 repos~7.4kAutomated safety check: NotesGPL-2.0
Advanced Evaluationaiskillstore/marketplace4304 repos~4.2kAutomated safety check: PassNone
Design AI BenchmarkingAperivue/medsci-skills329—~2.4kAutomated safety check: PassMIT
Quality Reportindranilbanerjee/digital-marketing-pro8541 repos~2.5kAutomated safety check: PassMIT
AI Eval Planmohitagw15856/pm-claude-skills1.4k—~996Automated safety check: PassMIT

Similar skills

  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub starsUsed in 1 repo~7.4k tokens
    EducationAuto-check: notes
  • Advanced Evaluation

    aiskillstore/marketplace

    This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

    430 GitHub starsUsed in 4 repos~4.2k tokens
    EducationAuto-check passed
  • Design AI Benchmarking

    Aperivue/medsci-skills

    A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

    329 GitHub stars~2.4k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Quality Report

    indranilbanerjee/digital-marketing-pro

    Report content-quality trends over time from logged evaluations: weekly score trend charts, a content-type leaderboard, per-dimension performance breakdown, statistically flagged regression alerts…

    854 GitHub starsUsed in 1 repo~2.5k tokens
    EducationAuto-check passed
  • AI Eval Plan

    mohitagw15856/pm-claude-skills

    Design an evaluation plan for an LLM or AI feature before shipping it.

    1.4k GitHub stars~996 tokensUpdated yesterday
    EducationAuto-check passed
  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    353 GitHub stars~4.5k tokensUpdated 2 days ago
    EducationAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Agentic Evaluation Framework

What does Agentic Evaluation Framework do?

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality". Agentic Evaluation Framework is an agent skill from borghei/Claude-Skills. This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

When should I use Agentic Evaluation Framework?

Agentic Evaluation Framework fits situations like: asks to evaluate LLM output quality; set up LLM-as-judge; build an eval rubric; compare model outputs pairwise.

How do I install Agentic Evaluation Framework in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill agentic-evaluation-framework -a claude-code`. Or copy the skill folder (engineering/agentic-evaluation-framework in borghei/Claude-Skills) into .claude/skills/agentic-evaluation-framework in your project. Claude Code loads it when a task matches its description.

How do I install Agentic Evaluation Framework in Codex?

Run `npx skills add borghei/Claude-Skills --skill agentic-evaluation-framework -a codex`. Or copy the skill folder (engineering/agentic-evaluation-framework in borghei/Claude-Skills) into .agents/skills/agentic-evaluation-framework in your project. Codex loads it when a task matches its description.

Can I use Agentic Evaluation Framework in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill agentic-evaluation-framework -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-evaluation-framework, .gemini/skills/agentic-evaluation-framework, .github/skills/agentic-evaluation-framework and .opencode/skills/agentic-evaluation-framework in your project.

What does Agentic Evaluation Framework need to run?

Going by SKILL.md and its folder, Agentic Evaluation Framework needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Agentic Evaluation Framework access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic Evaluation Framework safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agentic Evaluation Framework use?

Agentic Evaluation Framework is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic Evaluation Framework use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Agentic Evaluation Framework?

Skills that share tags, products or a category with Agentic Evaluation Framework: Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars), Advanced Evaluation (aiskillstore/marketplace, 430 stars), Design AI Benchmarking (Aperivue/medsci-skills, 329 stars) and Quality Report (indranilbanerjee/digital-marketing-pro, 854 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic Evaluation Framework?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.