Official agent skill

Agentic Eval

by github in github/awesome-copilot

Patterns and techniques for evaluating and improving AI agent outputs.

OfficialMITAuto-check passedAI & LLM Engineering

Install Agentic Eval

skills CLI
$ npx skills add github/awesome-copilot --skill agentic-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot agentic-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentic-eval .claude/skills/agentic-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-eval
GitHub stars
40k
Used in
4 other repos
Token cost
~1.5k tokens
SKILL.md length
185 words
Files
1
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Patterns and techniques for evaluating and improving AI agent outputs.

  • LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code
  • SKILL.md covers Overview, When to Use, Pattern 1: Basic Reflection and Pattern 2: Evaluator-Optimizer, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Analysis) - Measuring and improving agent response quality

What it does

Agentic Eval is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimizer pipelines for quality-critical generation - Creating test-driven code refinement workflows - Designing rubric-based or LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code, reports, analysis) - Measuring and improving agent response quality

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation, Quizzes and assessments and Test-driven development. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code
  • Analysis) - Measuring and improving agent response quality

Example prompts

  • “Use the agentic-eval skill to pattern and techniques for evaluating and improving AI agent outputs”
  • “/agentic-eval”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic Eval loads about 1.5k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 185 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 185 words, ~1,466 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-eval/SKILL.md (or your agent's skills folder).
name
agentic-eval
description
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimizer pipelines for quality-critical generation - Creating test-driven code refinement workflows - Designing rubric-based or LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code, reports, analysis) - Measuring and improving agent response quality

Agentic Evaluation Patterns

Patterns for self-improvement through iterative evaluation and refinement.

Overview

Evaluation patterns enable agents to assess and improve their own outputs, moving beyond single-shot generation to iterative refinement loops.

Generate → Evaluate → Critique → Refine → Output
    ↑                              │
    └──────────────────────────────┘

When to Use

  • Quality-critical generation: Code, reports, analysis requiring high accuracy
  • Tasks with clear evaluation criteria: Defined success metrics exist
  • Content requiring specific standards: Style guides, compliance, formatting

Pattern 1: Basic Reflection

Agent evaluates and improves its own output through self-critique.

python
def reflect_and_refine(task: str, criteria: list[str], max_iterations: int = 3) -> str:
    """Generate with reflection loop."""
    output = llm(f"Complete this task:\n{task}")
    
    for i in range(max_iterations):
        # Self-critique
        critique = llm(f"""
        Evaluate this output against criteria: {criteria}
        Output: {output}
        Rate each: PASS/FAIL with feedback as JSON.
        """)
        
        critique_data = json.loads(critique)
        all_pass = all(c["status"] == "PASS" for c in critique_data.values())
        if all_pass:
            return output
        
        # Refine based on critique
        failed = {k: v["feedback"] for k, v in critique_data.items() if v["status"] == "FAIL"}
        output = llm(f"Improve to address: {failed}\nOriginal: {output}")
    
    return output

Key insight: Use structured JSON output for reliable parsing of critique results.


Pattern 2: Evaluator-Optimizer

Separate generation and evaluation into distinct components for clearer responsibilities.

python
class EvaluatorOptimizer:
    def __init__(self, score_threshold: float = 0.8):
        self.score_threshold = score_threshold
    
    def generate(self, task: str) -> str:
        return llm(f"Complete: {task}")
    
    def evaluate(self, output: str, task: str) -> dict:
        return json.loads(llm(f"""
        Evaluate output for task: {task}
        Output: {output}
        Return JSON: {{"overall_score": 0-1, "dimensions": {{"accuracy": ..., "clarity": ...}}}}
        """))
    
    def optimize(self, output: str, feedback: dict) -> str:
        return llm(f"Improve based on feedback: {feedback}\nOutput: {output}")
    
    def run(self, task: str, max_iterations: int = 3) -> str:
        output = self.generate(task)
        for _ in range(max_iterations):
            evaluation = self.evaluate(output, task)
            if evaluation["overall_score"] >= self.score_threshold:
                break
            output = self.optimize(output, evaluation)
        return output

Pattern 3: Code-Specific Reflection

Test-driven refinement loop for code generation.

python
class CodeReflector:
    def reflect_and_fix(self, spec: str, max_iterations: int = 3) -> str:
        code = llm(f"Write Python code for: {spec}")
        tests = llm(f"Generate pytest tests for: {spec}\nCode: {code}")
        
        for _ in range(max_iterations):
            result = run_tests(code, tests)
            if result["success"]:
                return code
            code = llm(f"Fix error: {result['error']}\nCode: {code}")
        return code

Evaluation Strategies

Outcome-Based

Evaluate whether output achieves the expected result.

python
def evaluate_outcome(task: str, output: str, expected: str) -> str:
    return llm(f"Does output achieve expected outcome? Task: {task}, Expected: {expected}, Output: {output}")
LLM-as-Judge

Use LLM to compare and rank outputs.

python
def llm_judge(output_a: str, output_b: str, criteria: str) -> str:
    return llm(f"Compare outputs A and B for {criteria}. Which is better and why?")
Rubric-Based

Score outputs against weighted dimensions.

python
RUBRIC = {
    "accuracy": {"weight": 0.4},
    "clarity": {"weight": 0.3},
    "completeness": {"weight": 0.3}
}

def evaluate_with_rubric(output: str, rubric: dict) -> float:
    scores = json.loads(llm(f"Rate 1-5 for each dimension: {list(rubric.keys())}\nOutput: {output}"))
    return sum(scores[d] * rubric[d]["weight"] for d in rubric) / 5

Best Practices

PracticeRationale
Clear criteriaDefine specific, measurable evaluation criteria upfront
Iteration limitsSet max iterations (3-5) to prevent infinite loops
Convergence checkStop if output score isn't improving between iterations
Log historyKeep full trajectory for debugging and analysis
Structured outputUse JSON for reliable parsing of evaluation results

Quick Start Checklist

markdown
## Evaluation Implementation Checklist

### Setup
- [ ] Define evaluation criteria/rubric
- [ ] Set score threshold for "good enough"
- [ ] Configure max iterations (default: 3)

### Implementation
- [ ] Implement generate() function
- [ ] Implement evaluate() function with structured output
- [ ] Implement optimize() function
- [ ] Wire up the refinement loop

### Safety
- [ ] Add convergence detection
- [ ] Log all iterations for debugging
- [ ] Handle evaluation parse failures gracefully

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agentic-eval of github/awesome-copilot.

Open the folder on GitHubat commit 727ff2e

Used in 4 other repositories

We found 12 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agentic Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic Eval this skillgithub/awesome-copilot40k4 repos~1.5kAutomated safety check: PassMIT
Advanced Evaluationguanyang/open-agent-hub9732 repos~4.2kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT
Agentsop Metric Designagentsope/SkillAlchemy457—~6.4kAutomated safety check: PassMIT
Suede AI EvalJasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMIT
Commerce Evalsanthropics/commerce-agents3.2k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    973 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Metric Design

    agentsope/SkillAlchemy

    Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

    457 GitHub stars~6.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Suede AI Eval

    JasonColapietro/suede-creator-skills

    Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

    127 GitHub stars~3.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Commerce Evals

    anthropics/commerce-agents

    Official

    Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

    3.2k GitHub stars~1.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub starsUsed in 1 repo~7.4k tokens
    EducationAuto-check: notes

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Agentic Eval

What does Agentic Eval do?

Patterns and techniques for evaluating and improving AI agent outputs. Agentic Eval is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Patterns and techniques for evaluating and improving AI agent outputs.

When should I use Agentic Eval?

Agentic Eval fits situations like: LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code; analysis) - Measuring and improving agent response quality.

How do I install Agentic Eval in Claude Code?

Run `npx skills add github/awesome-copilot --skill agentic-eval -a claude-code`. Or copy the skill folder (skills/agentic-eval in github/awesome-copilot) into .claude/skills/agentic-eval in your project. Claude Code loads it when a task matches its description.

How do I install Agentic Eval in Codex?

Run `npx skills add github/awesome-copilot --skill agentic-eval -a codex`. Or copy the skill folder (skills/agentic-eval in github/awesome-copilot) into .agents/skills/agentic-eval in your project. Codex loads it when a task matches its description.

Can I use Agentic Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill agentic-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-eval, .gemini/skills/agentic-eval, .github/skills/agentic-eval and .opencode/skills/agentic-eval in your project.

What does Agentic Eval need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentic Eval is instructions for the agent only. Our summary lists: Python 3.

Does Agentic Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentic Eval use?

Agentic Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic Eval use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agentic Eval?

Skills that share tags, products or a category with Agentic Eval: Advanced Evaluation (guanyang/open-agent-hub, 973 stars), Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars), Agentsop Metric Design (agentsope/SkillAlchemy, 457 stars) and Suede AI Eval (JasonColapietro/suede-creator-skills, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic Eval?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.