Agent skill

Evaluation

by guanyang in guanyang/open-agent-hub

This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

MITAuto-check passedEducation

Install Evaluation

skills CLI
$ npx skills add guanyang/open-agent-hub --skill evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guanyang/open-agent-hub evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluation .claude/skills/evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation
GitHub stars
973
Used in
2 other repos
Token cost
~4.2k tokens
SKILL.md length
1,900 words
Files
3 (incl. scripts, references)
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

  • Works in 9 steps: Define quality dimensions relevant to… → Create rubrics with clear, descriptive… → Build test sets from real usage patterns… → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers When to Activate, Core Concepts, Detailed Topics and Practical Guidance, plus 6 more sections
  • Runs Python scripts from its folder

What it does

Evaluation is an agent skill from guanyang/open-agent-hub. This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/metrics.md` and `scripts/evaluator.py`).

It sits in Education, covering Quizzes and assessments, Quality gates and Building AI agents. The repository describes itself as: A lightweight, zero-dependency CLI tool to manage and activate capabilities for AI coding assistants (such as Claude Code, Cursor, Trae, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Quizzes and assessments
  • Tasks that involve Quality gates
  • Tasks that involve Building AI agents

Example prompts

  • “/evaluation”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Define quality dimensions relevant to the use case before writing any evaluation code, because dimensions chosen later tend to reflect…
  2. Create rubrics with clear, descriptive level definitions so evaluators (human or LLM) produce consistent scores.
  3. Build test sets from real usage patterns and edge cases, stratified by complexity, with at least 50 cases for reliable signal.
  4. Implement automated evaluation pipelines that run on every significant change.
  5. Establish baseline metrics before making changes so improvements can be measured against a known reference.
  6. Run evaluations on all significant changes and compare against the baseline.
  7. Track metrics over time for trend analysis because gradual degradation is harder to notice than sudden drops.
  8. Supplement automated evaluation with human review on a regular cadence.
  9. Separate deterministic validation failures from quality judgments so invalid artifacts cannot be laundered by a favorable LLM score.

What it can do on your machine

Read from SKILL.md and the folder at commit c32921b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation loads about 4.2k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 1,900 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from guanyang/open-agent-hub at commit c32921b, republished under its MIT licence (© guanyang). 1,900 words, ~4,208 tokens.

Download SKILL.mdSave it as .claude/skills/evaluation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evaluation
description
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.

Evaluation Methods for Agent Systems

Evaluate agent systems differently from traditional software because agents make dynamic decisions, are non-deterministic between runs, and often lack single correct answers. Build evaluation frameworks that account for these characteristics, provide actionable feedback, catch regressions, and validate that context engineering choices achieve intended effects.

When to Activate

Activate this skill when:

  • Testing agent performance systematically
  • Validating context engineering choices
  • Measuring improvements over time
  • Catching regressions before deployment
  • Building quality gates for agent pipelines
  • Comparing different agent configurations
  • Evaluating production systems continuously

Do not activate this skill for adjacent work owned by other skills:

  • Designing the LLM judge itself, pairwise comparison, judge calibration, or bias mitigation: advanced-evaluation.
  • Designing autonomous control surfaces, novelty gates, rollback, or PR approval boundaries: harness-engineering.
  • Debugging a specific context failure mode before measuring it: context-degradation.

Core Concepts

Focus evaluation on outcomes rather than execution paths, because agents may find alternative valid routes to goals. Judge whether the agent achieves the right outcome via a reasonable process, not whether it followed a specific sequence of steps.

Use multi-dimensional rubrics instead of single scores because one number hides critical failures in specific dimensions. Capture factual accuracy, completeness, citation accuracy, source quality, and tool efficiency as separate dimensions, then weight them for the use case.

Use model-judged evaluation only after deterministic checks and rubrics are stable. When the work centers on judge prompts, pairwise comparison, calibration, or bias mitigation, switch to Advanced Evaluation.

Run deterministic validation before LLM judgment whenever the artifact has machine-checkable structure. Schema validity, duplicate keys, rubric math, manifest sync, retrieval status, and required evidence paths should fail fast before an evaluator spends tokens or returns a subjective score.

Performance Drivers

Apply browsing-agent research when designing evaluation budgets: token usage, tool calls, and model choice can dominate measured performance variance (claim-evaluation-browsecomp-variance).

FactorVariance ExplainedImplication
Token usagePrimary driverMore exploration can improve performance until cost or context quality collapses
Number of tool callsSecondary driverMore tool use helps only when calls retrieve useful evidence
Model choiceSecondary but multiplicativeBetter models often use tokens and tools more efficiently

Act on these implications when designing evaluations:

  • Set realistic token budgets: Evaluate agents with production-realistic token limits, not unlimited resources.
  • Compare model upgrades against token increases: Better models may use tokens more efficiently than weaker models with larger budgets.
  • Validate multi-agent architectures: Extra agents add tokens and tool calls; evaluate them against single-agent baselines.

Detailed Topics

Evaluation Challenges

Handle Non-Determinism and Multiple Valid Paths

Design evaluations that tolerate path variation because agents may take completely different valid paths to reach goals. One agent might search three sources while another searches ten; both may produce correct answers. Avoid checking for specific steps. Instead, define outcome criteria (correctness, completeness, quality) and score against those, treating the execution path as informational rather than evaluative.

Test Context-Dependent Failures

Evaluate across a range of complexity levels and interaction lengths because agent failures often depend on context in subtle ways. An agent might succeed on simple queries but fail on complex ones, work well with one tool set but fail with another, or degrade after extended interaction as context accumulates. Include simple, medium, complex, and very complex test cases to surface these patterns.

Score Composite Quality Dimensions Separately

Break agent quality into separate dimensions (factual accuracy, completeness, coherence, tool efficiency, process quality) and score each independently because an agent might score high on accuracy but low on efficiency, or vice versa. Then compute weighted aggregates tuned to use-case priorities. This approach reveals which dimensions need improvement rather than averaging away the signal.

Evaluation Rubric Design

Build Multi-Dimensional Rubrics

Define rubrics covering key dimensions with descriptive levels from excellent to failed. Include these core dimensions and adapt weights per use case:

  • Factual accuracy: Claims match ground truth (weight heavily for knowledge tasks)
  • Completeness: Output covers requested aspects (weight heavily for research tasks)
  • Citation accuracy: Citations match claimed sources (weight for trust-sensitive contexts)
  • Source quality: Uses appropriate primary sources (weight for authoritative outputs)
  • Tool efficiency: Uses right tools a reasonable number of times (weight for cost-sensitive systems)

Convert Rubrics to Numeric Scores

Map dimension assessments to numeric scores (0.0 to 1.0), apply per-dimension weights, and calculate weighted overall scores. Set passing thresholds based on use-case requirements, typically 0.7 for general use and 0.9 for high-stakes applications. Store individual dimension scores alongside the aggregate because the breakdown drives targeted improvement.

Evaluation Methodologies

Use LLM-as-Judge for Scale

Build LLM-based evaluation prompts that include: clear task description, the agent output under test, ground truth when available, an evaluation scale with explicit level descriptions, and a request for structured judgment with reasoning. LLM judges provide consistent, scalable evaluation across large test sets. Use a different model family than the agent being evaluated to avoid self-enhancement bias.

Supplement with Human Evaluation

Route edge cases, unusual queries, and a random sample of production traffic to human reviewers because humans notice hallucinated answers, system failures, and subtle biases that automated evaluation misses. Track patterns across human reviews to identify systematic issues and feed findings back into automated evaluation criteria.

Apply End-State Evaluation for Stateful Agents

For agents that mutate persistent state (files, databases, configurations), evaluate whether the final state matches expectations rather than how the agent got there. Define expected end-state assertions and verify them programmatically after each test run.

Test Set Design

Select Representative Samples

Start with small samples (20-30 cases) during early development when changes have dramatic impacts and low-hanging fruit is abundant. Scale to 50+ cases for reliable signal as the system matures. Sample from real usage patterns, add known edge cases, and ensure coverage across complexity levels.

Stratify by Complexity

Structure test sets across complexity levels to prevent easy examples from inflating scores:

  • Simple: single tool call, factual lookup
  • Medium: multiple tool calls, comparison logic
  • Complex: many tool calls, significant ambiguity
  • Very complex: extended interaction, deep reasoning, synthesis

Report scores per stratum alongside overall scores to reveal where the agent actually struggles.

Context Engineering Evaluation

Validate Context Strategies Systematically

Run agents with different context strategies on the same test set and compare quality scores, token usage, and efficiency metrics. This isolates the effect of context engineering from other variables and prevents anecdote-driven decisions.

Run Degradation Tests

Test how context degradation affects performance by running agents at different context sizes. Identify performance cliffs where context becomes problematic and establish safe operating limits. Feed these limits back into context management strategies.

Continuous Evaluation

Build Automated Evaluation Pipelines

Integrate evaluation into the development workflow so evaluations run automatically on agent changes. Track results over time, compare versions, and block deployments that regress on key metrics.

Monitor Production Quality

Sample production interactions and evaluate them continuously. Set alerts for quality drops below warning (0.85 pass rate) and critical (0.70 pass rate) thresholds. Maintain dashboards showing trend analysis over time windows to detect gradual degradation.

Practical Guidance

Show full SKILL.md (761 more words)Show less
Building Evaluation Frameworks

Follow this sequence to build an evaluation framework, because skipping early steps leads to measurements that do not reflect real quality:

  1. Define quality dimensions relevant to the use case before writing any evaluation code, because dimensions chosen later tend to reflect what is easy to measure rather than what matters.
  2. Create rubrics with clear, descriptive level definitions so evaluators (human or LLM) produce consistent scores.
  3. Build test sets from real usage patterns and edge cases, stratified by complexity, with at least 50 cases for reliable signal.
  4. Implement automated evaluation pipelines that run on every significant change.
  5. Establish baseline metrics before making changes so improvements can be measured against a known reference.
  6. Run evaluations on all significant changes and compare against the baseline.
  7. Track metrics over time for trend analysis because gradual degradation is harder to notice than sudden drops.
  8. Supplement automated evaluation with human review on a regular cadence.
  9. Separate deterministic validation failures from quality judgments so invalid artifacts cannot be laundered by a favorable LLM score.
Avoiding Evaluation Pitfalls

Guard against these common failures that undermine evaluation reliability:

  • Overfitting to specific paths: Evaluate outcomes, not specific steps, because agents find novel valid paths.
  • Ignoring edge cases: Include diverse test scenarios covering the full complexity spectrum.
  • Single-metric obsession: Use multi-dimensional rubrics because a single score hides dimension-specific failures.
  • Neglecting context effects: Test with realistic context sizes and histories rather than clean-room conditions.
  • Skipping human evaluation: Automated evaluation misses subtle issues that humans catch reliably.

Examples

Example 1: Simple Evaluation

python
def evaluate_agent_response(response, expected):
    rubric = load_rubric()
    scores = {}
    for dimension, config in rubric.items():
        scores[dimension] = assess_dimension(response, expected, dimension)
    overall = weighted_average(scores, config["weights"])
    return {"passed": overall >= 0.7, "scores": scores}

Example 2: Test Set Structure

Test sets should span multiple complexity levels to ensure comprehensive evaluation:

python
test_set = [
    {
        "name": "simple_lookup",
        "input": "What is the capital of France?",
        "expected": {"type": "fact", "answer": "Paris"},
        "complexity": "simple",
        "description": "Single tool call, factual lookup"
    },
    {
        "name": "medium_query",
        "input": "Compare the revenue of Apple and Microsoft last quarter",
        "complexity": "medium",
        "description": "Multiple tool calls, comparison logic"
    },
    {
        "name": "multi_step_reasoning",
        "input": "Analyze sales data from Q1-Q4 and create a summary report with trends",
        "complexity": "complex",
        "description": "Many tool calls, aggregation, analysis"
    },
    {
        "name": "research_synthesis",
        "input": "Research emerging AI technologies, evaluate their potential impact, and recommend adoption strategy",
        "complexity": "very_complex",
        "description": "Extended interaction, deep reasoning, synthesis"
    }
]

Example 3: Deterministic gate before model judgment

python
def evaluate_pr_candidate(candidate):
    structure = run_validate_repo(candidate)
    if not structure.ok:
        return {"passed": False, "reason": "deterministic validation failed", "details": structure.errors}

    quality = run_rubric_eval(candidate)
    return {"passed": quality.overall >= 0.8, "scores": quality.dimensions}

Example 4: Quality gate dimensions

yaml
gate:
  deterministic:
    - schema_valid
    - required_files_present
    - no_duplicate_ids
  quality:
    factual_accuracy: min 0.85
    completeness: min 0.80
    source_traceability: min 0.90

Guidelines

  1. Use multi-dimensional rubrics, not single metrics
  2. Evaluate outcomes, not specific execution paths
  3. Cover complexity levels from simple to complex
  4. Test with realistic context sizes and histories
  5. Run evaluations continuously, not just before release
  6. Supplement LLM evaluation with human review
  7. Track metrics over time for trend detection
  8. Set clear pass/fail thresholds based on use case

Gotchas

  1. Overfitting evals to specific code paths: Tests pass but the agent fails on slight input variations. Write eval criteria against outcomes and semantics, not surface patterns, and rotate test inputs periodically.
  2. LLM-judge self-enhancement bias: Models rate their own outputs higher than independent judges do. Use a different model family as the evaluation judge than the model being evaluated.
  3. Test set contamination: Eval examples leak into training data or prompt templates, inflating scores. Keep eval sets versioned and separate from any data used in prompts or fine-tuning.
  4. Metric gaming: Optimizing for the metric rather than actual quality produces agents that score well but disappoint users. Cross-validate automated metrics against human judgments regularly.
  5. Single-dimension scoring: One aggregate number hides critical failures in specific dimensions. Always report per-dimension scores alongside the overall score, and fail the eval if any single dimension falls below its minimum threshold.
  6. Eval set too small: Fewer than 50 examples produces unreliable signal with high variance between runs. Scale the eval set to at least 50 cases and report confidence intervals.
  7. Not stratifying by difficulty: Easy examples inflate overall scores, masking failures on hard cases. Report scores per complexity stratum and weight the overall score to prevent easy-case dominance.
  8. Treating eval as one-time: Evaluation must be continuous, not a launch gate. Agent quality drifts as models update, tools change, and usage patterns evolve. Run evals on every change and on a regular production cadence.

Integration

This skill owns outcome measurement and quality gates. Adjacent skills own specialized evaluator design and control-loop governance:

  • advanced-evaluation: LLM-as-judge prompt design, pairwise comparison, calibration, and bias mitigation.
  • harness-engineering: locked evaluators, editable surfaces, rollback, and human approval boundaries.
  • context-degradation: detecting and measuring degradation patterns.
  • context-optimization: measuring token, cost, latency, and quality effects of optimizations.
  • multi-agent-patterns: evaluating coordination quality and parallelization trade-offs.
  • tool-design: evaluating tool selection and recovery effectiveness.
  • memory-systems: evaluating memory retrieval and retention quality.

References

Internal reference:

  • Metrics Reference - Read when: designing specific evaluation metrics, choosing scoring scales, or implementing weighted rubric calculations

Internal skills:

  • All other skills connect to evaluation for quality measurement

External resources:

  • LLM evaluation benchmarks - Read when: selecting or building benchmark suites for agent comparison
  • Agent evaluation research papers - Read when: adopting new evaluation methodologies or validating current approach
  • Production monitoring practices - Read when: setting up alerting, dashboards, or sampling strategies for live systems

Skill Metadata

Created: 2025-12-20 Last Updated: 2026-05-15 Author: Agent Skills for Context Engineering Contributors Version: 1.2.0

© guanyang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/evaluation of guanyang/open-agent-hub.

  • SKILL.md
  • references/metrics.md
  • scripts/evaluator.py

Open the folder on GitHubat commit c32921b

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in guanyang/open-agent-hub, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation this skillguanyang/open-agent-hub9732 repos~4.2kAutomated safety check: PassMIT
Update Parker Skillreal-simple-labs/parker-brain100—~4.6kAutomated safety check: PassCustom licence
Agentsop Agent Topology Selectionagentsope/SkillAlchemy457—~4.7kAutomated safety check: PassMIT
Evaluation Frameworkathola/claude-night-market342—~1.3kAutomated safety check: PassMIT
Prompt Generatorcatlog22/Claude-Code-Workflow2.1k1 repos~4.7kAutomated safety check: NotesMIT
01 Auto Arenaagentscope-ai/OpenJudge8671 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Update Parker Skill

    real-simple-labs/parker-brain

    Make a correct update to Parker's prompts, system docs, rubrics, knowledge docs, training corpus, or brand outputs — and propagate the change everywhere it needs to land.

    100 GitHub stars~4.6k tokensUpdated today
    EducationAuto-check passed
  • Agentsop Agent Topology Selection

    agentsope/SkillAlchemy

    Cross-framework enhancement overlay for choosing a multi-agent topology BEFORE writing any agent.

    457 GitHub stars~4.7k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Evaluation Framework

    athola/claude-night-market

    Provides weighted scoring, rubrics, and decision-threshold patterns.

    342 GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Prompt Generator

    catlog22/Claude-Code-Workflow

    Generate or convert Claude Code prompt files — command orchestrators, skill files, agent role definitions, or style conversion of existing files.

    2.1k GitHub starsUsed in 1 repo~4.7k tokens
    Testing & QAAuto-check: notes
  • 01 Auto Arena

    agentscope-ai/OpenJudge

    Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

    867 GitHub starsUsed in 1 repo~2.5k tokens
    EducationAuto-check passed
  • Review a FHIRPath implementation change in Pathling against a correctness rubric covering collection semantics, empty propagation, column cardinality, type coercion, error-vs-empty behaviour, spec…

    137 GitHub stars~2k tokensUpdated today
    EducationAuto-check passed

More from guanyang/open-agent-hub

All 26 skills in this repo
  • Context Compression

    guanyang/open-agent-hub

    This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve…

    973 GitHub starsUsed in 2 repos~4.6k tokens
    Auto-check passed
  • Context Fundamentals

    guanyang/open-agent-hub

    This skill should be used to explain or reason about the foundational concepts of context engineering: what context is, the anatomy of a context window, how attention mechanics work, the U-shaped…

    973 GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed
  • Multi Agent Patterns

    guanyang/open-agent-hub

    This skill should be used when designing multi-agent systems that need context isolation, supervisor or swarm coordination, explicit handoffs, parallel execution, or a decision on whether multiple…

    973 GitHub starsUsed in 2 repos~4.6k tokens
    Auto-check passed
  • Project Development

    guanyang/open-agent-hub

    This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…

    973 GitHub starsUsed in 2 repos~4.7k tokens
    Auto-check passed
  • Tool Design

    guanyang/open-agent-hub

    This skill should be used for the tool-interface layer of an agent system specifically: writing tool descriptions agents can route on, designing tool schemas and response formats, naming…

    973 GitHub starsUsed in 2 repos~5k tokens
    Auto-check passed
  • Filesystem Context

    guanyang/open-agent-hub

    This skill should be used when agent work needs file-backed context: durable scratchpads, tool-output offloading, just-in-time discovery, cross-agent handoff files, filesystem memory, or cleanup…

    973 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed

Questions about Evaluation

What does Evaluation do?

This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…. Evaluation is an agent skill from guanyang/open-agent-hub. This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.

When should I use Evaluation?

Evaluation fits situations like: tasks that involve Quizzes and assessments; tasks that involve Quality gates; tasks that involve Building AI agents.

How do I install Evaluation in Claude Code?

Run `npx skills add guanyang/open-agent-hub --skill evaluation -a claude-code`. Or copy the skill folder (skills/evaluation in guanyang/open-agent-hub) into .claude/skills/evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation in Codex?

Run `npx skills add guanyang/open-agent-hub --skill evaluation -a codex`. Or copy the skill folder (skills/evaluation in guanyang/open-agent-hub) into .agents/skills/evaluation in your project. Codex loads it when a task matches its description.

Can I use Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guanyang/open-agent-hub --skill evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation, .gemini/skills/evaluation, .github/skills/evaluation and .opencode/skills/evaluation in your project.

What does Evaluation need to run?

Going by SKILL.md and its folder, Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluation use?

Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluation use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Evaluation?

Skills that share tags, products or a category with Evaluation: Update Parker Skill (real-simple-labs/parker-brain, 100 stars), Agentsop Agent Topology Selection (agentsope/SkillAlchemy, 457 stars), Evaluation Framework (athola/claude-night-market, 342 stars) and Prompt Generator (catlog22/Claude-Code-Workflow, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation?

guanyang (a GitHub user) maintains it in guanyang/open-agent-hub, which has 973 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.

Source: guanyang/open-agent-hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.