---
name: ak-dev-new-evaluator-provider
description: >
  Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel
  (beyond DeepEval, Opik and JEV). Use this skill when you need to give the test framework's pluggable
  AKEvaluator interface a new first-party scoring/judge backend addressable by a short
  config name (e.g. "trulens"), not a one-off bring-your-own evaluator. Covers implementing
  score-based and LLM-as-judge evaluation, factory registration, configuration, optional
  dependencies, and testing.
license: Apache-2.0
metadata:
  author: yaalalabs
  category: developer
---

# Adding a New Test Evaluator Provider

This guide walks through adding a new **built-in** evaluator provider to Agent Kernel's test
framework. Use the existing DeepEval implementation
(`ak-py/src/agentkernel/test/core/evaluator/deepeval.py`) as reference.

Before starting, check whether you actually need this skill: if the evaluator only needs to exist
for *your own* project (not addressable by every AK user via a short built-in name), you don't
need any of the steps below — just subclass `AKEvaluator` anywhere importable and point
`test-config.yaml`'s `evaluator:` at its dotted path. That's the "bring your own evaluator"
path described in the user-facing `ak-test` skill and
[`docs/docs/testing/cli-testing.md`](../../../docs/docs/testing/cli-testing.md#bring-your-own-evaluator);
`examples/cli/custom-evaluator/` is a complete worked example of it. This skill is only for adding
a **first-party, in-repo** provider that ships with AK and gets its own short `type` name.

## Existing Providers

| Provider | Short name | Scoring mode | LLM-judge mode | Extra |
|---|---|---|---|---|
| DeepEval | `deepeval` | `Scorer.quasi_exact_match_score` (whole-string, normalised) | `GEval` LLM-as-judge metric | `agentkernel[test]` |
| Opik | `opik` | `LevenshteinRatio` (fuzzy string similarity) | `GEval` LLM-as-judge metric | `agentkernel[opik]` |
| JEV | `jev` | — (`AKMetricNotSupported`) | TypeSafe Noul (yes/no probability) | `agentkernel[jev]` |

## Architecture Overview

- **`AKEvaluator`** (`ak-py/src/agentkernel/test/core/evaluator/base.py`) is the abstract base
  every evaluator — built-in or bring-your-own — implements. It has exactly two abstract methods:
  - `evaluate_by_score(case: AKEvaluationCase) -> AKEvaluationResult` — deterministic scoring,
    no LLM call
  - `evaluate_by_llm(case: AKEvaluationCase) -> AKEvaluationResult` — LLM-as-judge scoring
- **`AKEvaluationCase`** carries the comparison inputs (`user_input`, `actual`, `expected`,
  `threshold`, `context`, `criteria`); **`AKEvaluationResult`** carries the outcome (`score`,
  `passed`, `metric`, `evaluator`, `reason`, `cost`, `attempts`, `metadata`).
- **Error contract** (every evaluator must honor this, not just DeepEval):
  - Raise `AKMissingInput` when a field the requested metric needs (e.g. `case.expected`) wasn't
    supplied.
  - Raise `AKMetricNotSupported` from whichever of the two methods your backend structurally
    cannot implement (e.g. a pure LLM-judge service has no offline scoring mode).
  - Raise `AKEvaluationError` when a configured backend fails to produce a score (missing
    credentials, transport error, unparseable judge output). Never raise `AssertionError` and
    never silently return a `0.0` to stand in for a failure — `0.0` must only ever mean "scored
    zero", not "couldn't be scored". `Test.compare` is the only place that decides pass/fail
    fatality; evaluators only ever set `result.passed`.
- **`Test._resolve_evaluator_class`** in `ak-py/src/agentkernel/test/test.py` is the factory. It
  shares the same pluggable-backend shape as guardrails, sandbox providers, and trace backends
  (`core/util/factory.py`'s `resolve_dotted`/`require_extra`/`AKConfigError`): an `if`-per-built-in
  branch with the SDK import wrapped in `require_extra` (actionable `ImportError` naming the pip
  extra if missing), then a dotted-path bring-your-own fallback for anything else.
- Evaluator instances are cached per-process on `Test._evaluator` (keyed by the configured value),
  guarded by `Test._evaluator_lock` — construction happens once per distinct `evaluator:` config
  value, not once per `Test.compare` call.

## Step-by-Step

### 1. Create the Evaluator Provider File

Create `ak-py/src/agentkernel/test/core/evaluator/<provider>.py`. Keep the provider's SDK imports
inside this file only — `test/core/evaluator/__init__.py` and `base.py` stay pure Python with no
optional-dependency imports at module level, so importing the `AKEvaluator` interface never
requires your provider's SDK to be installed.

```python
# ak-py/src/agentkernel/test/core/evaluator/<provider>.py
from agentkernel.test.config import AKTestConfig

from .base import AKEvaluationCase, AKEvaluationError, AKEvaluationResult, AKEvaluator, AKMissingInput


class <Provider>AKEvaluator(AKEvaluator):
    def __init__(self, config: AKTestConfig) -> None:
        super().__init__(config)
        # Lazy-init any client/model here only if evaluate_by_score never needs it
        # (mirrors DeepevalAKEvaluator's lazy LiteLLMModel, built only on first evaluate_by_llm call).

    def evaluate_by_score(self, case: AKEvaluationCase) -> AKEvaluationResult:
        if not case.expected:
            raise AKMissingInput("evaluate_by_score requires AKEvaluationCase.expected")
        # Deterministic, offline scoring logic here.
        score = ...  # float
        return AKEvaluationResult(
            metric="<metric_name>",
            evaluator="<provider>",
            score=score,
            passed=score >= case.threshold,
        )

    def evaluate_by_llm(self, case: AKEvaluationCase) -> AKEvaluationResult:
        if not case.expected:
            raise AKMissingInput("evaluate_by_llm requires AKEvaluationCase.expected")
        try:
            score = ...  # call the judge
        except Exception as exc:
            raise AKEvaluationError(f"<provider> llm-based evaluation failed: {exc}") from exc
        return AKEvaluationResult(
            metric="<metric_name>",
            evaluator="<provider>",
            score=score,
            reason=...,  # judge's explanation, if the backend provides one
            passed=score is not None and score >= case.threshold,
        )
```

If a mode genuinely doesn't apply to your backend (e.g. a provider that is LLM-judge-only), raise
`AKMetricNotSupported` from that method instead of faking a result. `Test.compare` does not catch
it — in `fallback` mode it propagates out of `evaluate_by_score` before `evaluate_by_llm` runs — so
document that users of your provider must set the matching `mode` (e.g. JEV requires `mode: llm`).

### 2. Register with the Factory

Add the short name to `_BUILTIN_EVALUATORS` and a branch in `Test._resolve_evaluator_class`, both
in `ak-py/src/agentkernel/test/test.py`:

```python
_BUILTIN_EVALUATORS = ["deepeval", "opik", "jev", "<provider>"]          # ADD THIS

class Test:
    ...
    @classmethod
    def _resolve_evaluator_class(cls, configured: str) -> type[AKEvaluator]:
        if configured == "deepeval":
            with require_extra("test", "evaluator: deepeval"):
                from .core.evaluator.deepeval import DeepevalAKEvaluator
            return DeepevalAKEvaluator
        if configured == "opik":
            with require_extra("opik", "evaluator: opik"):
                from .core.evaluator.opik import OpikAKEvaluator
            return OpikAKEvaluator
        if configured == "jev":
            with require_extra("jev", "evaluator: jev"):
                from .core.evaluator.jev import JevAKEvaluator
            return JevAKEvaluator
        if configured == "<provider>":                                        # ADD THIS
            with require_extra("<provider>", "evaluator: <provider>"):
                from .core.evaluator.<provider> import <Provider>AKEvaluator
            return <Provider>AKEvaluator
        if "." not in configured:
            raise AKConfigError(
                f"unknown evaluator '{configured}'; expected one of {_BUILTIN_EVALUATORS} or a dotted path to an AKEvaluator subclass"
            )
        return resolve_dotted(configured, base=AKEvaluator)
```

A dotted `evaluator:` value (e.g. `myorg.evaluators.CustomEvaluator`) resolves via `resolve_dotted`
without any factory edit at all — only add an `if` branch here for a first-party, in-repo provider
you want addressable by a short name.

### 3. Add Optional Dependencies

Add a new extras group to `ak-py/pyproject.toml` for the provider's SDK — don't fold it into the
existing `test` extra (that one stays DeepEval's, since every test user already needs it for the
framework itself). Follow the pattern of the `opik` extra, the first provider added on top of the
original DeepEval-only `test` extra:

```toml
[project.optional-dependencies]
<provider> = [
    "provider-sdk>=x.y.z",
]
```

### 4. Add Configuration Docs

`evaluator:` in `test-config.yaml` is already a free-form string on `AKTestConfig` (built-in short
name or dotted path) — no config schema change is needed for a new built-in, since it's just a new
value the same field accepts:

```yaml
mode: fallback
evaluator: <provider>
```

If your provider needs extra config fields (e.g. an API key env var name, a judge model override),
read them from `AKTestConfig` the same way `DeepevalAKEvaluator` reads `self._config.llm` — don't
invent a parallel config path.

### 5. Add Tests

Add `ak-py/tests/test_evaluator_<provider>.py`, following the shape of
`ak-py/tests/test_evaluator_deepeval.py`: exercise `evaluate_by_score` for real (offline, no
network) where possible, and mock the judge call in `evaluate_by_llm` so the suite stays
network-free. At minimum cover:

- `evaluate_by_score`: exact/mismatch cases, threshold boundary, `AKMissingInput` when `expected`
  is absent
- `evaluate_by_llm`: success, failure wrapped as `AKEvaluationError`, `AKMissingInput` when
  `expected` is absent
- The factory branch: `Test._resolve_evaluator_class("<provider>")` resolves to your class, and
  (if the SDK is optional) the `require_extra` `ImportError` path when it's missing — see
  `test_resolve_evaluator_class_deepeval_missing_extra_raises_import_error` in
  `ak-py/tests/test_cli_tester.py` for the pattern (patching `builtins.__import__`, since a
  cached submodule import can otherwise mask the missing dependency).

### 6. Add an Example

Add `examples/cli/<provider>-evaluator/`, following the shape of `examples/cli/opik-evaluator/`
(a minimal agent, a `demo_test.py` exercising the new evaluator, and a `test-config.yaml` pointing
`evaluator:` at the new short name). Register it in `.github/test-config.yaml`'s e2e matrix so it
runs in CI, the way every other `examples/cli/*` entry does.

### 7. Add Documentation

Neither doc page carries a literal "evaluator backend table" — both describe the built-ins in
prose next to the `score`/`llm`/`fallback` mode explanations. Update every prose mention that
enumerates the built-ins by name, not just one page:

- [`docs/docs/core-concepts/configuration.md`](../../../docs/docs/core-concepts/configuration.md)
  and [`docs/docs/testing/cli-testing.md`](../../../docs/docs/testing/cli-testing.md) — the
  `evaluator:` field description and the score/llm mode explanations.
- [`docs/docs/testing/automated-testing.md`](../../../docs/docs/testing/automated-testing.md) and
  [`docs/docs/testing/overview.md`](../../../docs/docs/testing/overview.md) — same prose pattern,
  duplicated across these pages.
- [`docs/docs/agent-skills.md`](../../../docs/docs/agent-skills.md) — the skill directory rows for
  this skill and for `ak-dev-testing-conventions`.
- `.agents/skills/ak-dev-testing-conventions/SKILL.md` — the evaluator config/mode section.
- `ak-py/README.md` — the Test Configuration reference (`evaluator` field) and the test-config
  walkthrough section.
- The user-facing `ak-test` skill (`ak-py/src/agentkernel/skills/ak-test/SKILL.md`) and its
  `evals/evals.json`.
- Landing page inventories (`docs/src/components/*/data.tsx`): a tile in the **Observability,
  safety & testing** row of `IntegrationsMarquee/data.tsx` (role `Evaluator`, `href` to the
  automated testing page, logo or `react-icons/si` glyph), and the provider in the **Pluggable
  Evaluators** card's `tags` and `description` under the Observe tab in
  `FeatureExplorer/data.tsx`. Logo sourcing and the build check are in
  `ak-dev-sync-docs-from-branch`, *Docs-Site Landing and Features Pages*.
- The features page (`docs/src/pages/features.tsx`): the `approaches` entry for Pluggable
  Evaluators under Testing & Evaluation names every built-in.

## Checklist

- [ ] `ak-py/src/agentkernel/test/core/evaluator/<provider>.py` implementing `AKEvaluator`
- [ ] Factory registration in `Test._resolve_evaluator_class` (`ak-py/src/agentkernel/test/test.py`)
      and `_BUILTIN_EVALUATORS`
- [ ] Optional dependency extra in `ak-py/pyproject.toml`
- [ ] Unit tests in `ak-py/tests/test_evaluator_<provider>.py`
- [ ] Example in `examples/cli/<provider>-evaluator/`, registered in `.github/test-config.yaml`
- [ ] Documentation updated: `docs/docs/core-concepts/configuration.md`,
      `docs/docs/testing/cli-testing.md`, `docs/docs/testing/automated-testing.md`,
      `docs/docs/testing/overview.md`, `docs/docs/agent-skills.md`,
      `.agents/skills/ak-dev-testing-conventions/SKILL.md`, `ak-py/README.md`, the `ak-test` skill
      and its `evals/evals.json`
- [ ] Landing page inventories: marquee tile (`IntegrationsMarquee/data.tsx`), Pluggable Evaluators card
      tags (`FeatureExplorer/data.tsx`); the Pluggable Evaluators `approaches` entry in `features.tsx`
