Agent skill

Add Evaluator

by wso2 in wso2/agent-manager

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Evaluator

skills CLI
$ npx skills add wso2/agent-manager --skill add-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wso2/agent-manager add-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wso2/agent-manager.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-evaluator .claude/skills/add-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-evaluator
GitHub stars
105
Token cost
~710 tokens
SKILL.md length
228 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

  • Works in 7 steps: Pick the file — a built-in goes in… → Write the function and decorate it → Choose the level via the first… → …
  • The user asks to add
  • SKILL.md covers Steps, Gotchas, Commands (from… and Done checklist
  • Calls ruff, pip and pytest

What it does

Add Evaluator is an agent skill from wso2/agent-manager. Add a new evaluator to the amp-evaluation Python library. Use when the user asks to add, write, or register an evaluator, LLM-as-judge, or scoring check for agent traces in libs/amp-evaluation. Covers the decorator API and the type-hint-driven level/mode detection that determines whether the evaluator runs at trace/agent/LLM level and in experiment vs monitor mode.

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Type safety and LLM evaluation. It works with Python. The repository describes itself as: WSO2 AI Agent Manager is an open control plane designed for enterprises to deploy, manage, and govern AI agents at scale. The licence is Apache-2.0.

When your agent uses it

  • The user asks to add
  • Register an evaluator
  • Scoring check for agent traces in libs/amp-evaluation

Example prompts

  • “/add-evaluator”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Pick the file — a built-in goes in src/amp_evaluation/evaluators/builtin/ (standard.py for rule-based, llm_judge.py for judges…
  2. Write the function and decorate it
  3. Choose the level via the first parameter's type hint: Trace → TRACE, AgentTrace → AGENT, LLMSpan → LLM.
  4. Choose the mode via the task parameter
  5. LLM-as-judge: implement build_prompt() (not evaluate()) with the same level/mode detection; tag ["llm-judge", ]. Needs LLM config via the…
  6. Expose config knobs with the Param descriptor: max_latency_ms: float = Param(default=5000, description="…").
  7. Name must be unique — collisions are rejected in runner.run().

What it can do on your machine

Read from SKILL.md and the folder at commit 6d4af26. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ruff
    • pip
    • pytest
    • black
    • mypy

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Evaluator loads about 710 tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 228 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~710

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wso2/agent-manager at commit 6d4af26, republished under its Apache-2.0 licence (© wso2). 228 words, ~710 tokens.

Download SKILL.mdSave it as .claude/skills/add-evaluator/SKILL.md (or your agent's skills folder).
name
add-evaluator
description
Add a new evaluator to the amp-evaluation Python library. Use when the user asks to add, write, or register an evaluator, LLM-as-judge, or scoring check for agent traces in libs/amp-evaluation. Covers the decorator API and the type-hint-driven level/mode detection that determines whether the evaluator runs at trace/agent/LLM level and in experiment vs monitor mode.

Add an evaluator (amp-evaluation)

Read first: libs/amp-evaluation/AGENTS.md → "Defining an evaluator". This skill is the executable checklist. The non-obvious part is that level and mode are inferred from type hints, not declared.

Steps

  1. Pick the file — a built-in goes in src/amp_evaluation/evaluators/builtin/ (standard.py for rule-based, llm_judge.py for judges, deepeval.py for DeepEval wrappers). A user-defined one can live anywhere and be picked up by discover_evaluators(module).
  2. Write the function and decorate it:
    python
    from amp_evaluation import evaluator, Trace, Task, EvalResult
    
    @evaluator("my-check", description="…", tags=["rule-based", "quality"])
    def evaluate(trace: Trace) -> EvalResult:
        ...
  3. Choose the level via the first parameter's type hint: Trace → TRACE, AgentTrace → AGENT, LLMSpan → LLM.
  4. Choose the mode via the task parameter:
    • required task: Task → EXPERIMENT only;
    • task: Optional[Task] = None → both experiment and monitor;
    • no task param → both.
  5. LLM-as-judge: implement build_prompt() (not evaluate()) with the same level/mode detection; tag ["llm-judge", <aspect>]. Needs LLM config via the any-llm extra.
  6. Expose config knobs with the Param descriptor: max_latency_ms: float = Param(default=5000, description="…").
  7. Name must be unique — collisions are rejected in runner.run().

Gotchas

  • Type-hint detection uses typing.get_type_hints() — keep annotations importable (avoid forward refs that can't resolve).
  • semantic_similarity-style judges need expected_output on the task, so they're EXPERIMENT-only by nature.
  • Return an EvalResult; aggregations (mean/stddev) are computed per-evaluator by the runner from its scores.

Commands (from libs/amp-evaluation/)

bash
pip install -e '.[dev]'
pytest                          # runs with coverage (see pyproject)
ruff check src/                 # lint (line-length 120)
black src/                      # format (project configures Black; don't also run `ruff format`)
mypy src/                       # type-check

Done checklist

  • Correct level from the first-arg type hint; correct mode from the task param.
  • Unique name; tags set (rule-based / llm-judge + aspect).
  • Test added under tests/; pytest passes.
  • ruff check + mypy clean.

© wso2, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/add-evaluator of wso2/agent-manager.

Open the folder on GitHubat commit 6d4af26

Compare with similar skills

Add Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Evaluator this skillwso2/agent-manager105—~710Automated safety check: PassApache-2.0
Pydanticaimagnus919/agent-skills113—~4kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Formattingbrendanhasz/probflow175—~381Automated safety check: PassMIT
Pydantic AIdavila7/claude-code-templates32k4 repos~2.9kAutomated safety check: PassMIT
Pydantic AIdiegosouzapw/awesome-omni-skills159—~3.2kAutomated safety check: PassMIT

Similar skills

  • Pydanticai

    magnus919/agent-skills

    Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph.

    113 GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Formatting

    brendanhasz/probflow

    Ensure consistent code formatting using the uv package manager and pre-commit.

    175 GitHub stars~381 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Pydantic AI

    davila7/claude-code-templates

    Build production-ready AI agents with PydanticAI — type-safe tool use, structured outputs, dependency injection, and multi-model support.

    32k GitHub starsUsed in 4 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Pydantic AI

    diegosouzapw/awesome-omni-skills

    PydanticAI — Typed AI Agents in Python workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from wso2/agent-manager

  • Add Audit Event

    wso2/agent-manager

    Add a semantic audit event to agent-manager-service (the Go control plane).

    105 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Add Console API Feature

    wso2/agent-manager

    Add an API-backed feature to the console (React/TypeScript web UI).

    105 GitHub stars~761 tokensUpdated 2 days ago
    Auto-check passed
  • Add Service Unit Test

    wso2/agent-manager

    Write a service-layer unit test in agent-manager-service (the Go control plane).

    105 GitHub stars~997 tokensUpdated 2 days ago
    Auto-check passed
  • Add API Resource

    wso2/agent-manager

    Add or change a REST API resource in agent-manager-service (the Go control plane).

    105 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Add Evaluator

What does Add Evaluator do?

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager. Add Evaluator is an agent skill from wso2/agent-manager. Add a new evaluator to the amp-evaluation Python library.

When should I use Add Evaluator?

Add Evaluator fits situations like: the user asks to add; register an evaluator; scoring check for agent traces in libs/amp-evaluation.

How do I install Add Evaluator in Claude Code?

Run `npx skills add wso2/agent-manager --skill add-evaluator -a claude-code`. Or copy the skill folder (.claude/skills/add-evaluator in wso2/agent-manager) into .claude/skills/add-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Add Evaluator in Codex?

Run `npx skills add wso2/agent-manager --skill add-evaluator -a codex`. Or copy the skill folder (.claude/skills/add-evaluator in wso2/agent-manager) into .agents/skills/add-evaluator in your project. Codex loads it when a task matches its description.

Can I use Add Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wso2/agent-manager --skill add-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-evaluator, .gemini/skills/add-evaluator, .github/skills/add-evaluator and .opencode/skills/add-evaluator in your project.

What does Add Evaluator need to run?

Going by SKILL.md and its folder, Add Evaluator needs the command-line tools its instructions call (ruff, pip, pytest, black and mypy). Our summary lists: Python 3.

Does Add Evaluator access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Evaluator use?

Add Evaluator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Evaluator use?

About 710 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Evaluator?

Skills that share tags, products or a category with Add Evaluator: Pydanticai (magnus919/agent-skills, 113 stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Formatting (brendanhasz/probflow, 175 stars) and Pydantic AI (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Evaluator?

wso2 (a GitHub organization) maintains it in wso2/agent-manager, which has 105 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: wso2/agent-manager on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.