Agent skill

Ak Dev New Evaluator Provider

by yaalalabs in yaalalabs/agent-kernel

Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Ak Dev New Evaluator Provider

skills CLI
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .claude/skills/ak-dev-new-evaluator-provider && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ak-dev-new-evaluator-provider
GitHub stars
192
Token cost
~3.4k tokens
SKILL.md length
1,054 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

  • Works in 7 steps: Create the Evaluator Provider File → Register with the Factory → Add Optional Dependencies → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Existing Providers, Architecture Overview, Step-by-Step and Checklist
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ak Dev New Evaluator Provider is an agent skill from yaalalabs/agent-kernel. Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Use this skill when you need to give the test framework's pluggable AKEvaluator interface a new first-party scoring/judge backend addressable by a short config name (e.g. "trulens"), not a one-off bring-your-own evaluator. Covers implementing score-based and LLM-as-judge evaluation, factory registration, configuration, optional dependencies, and testing.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “trulens”
  • “/ak-dev-new-evaluator-provider”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Create the Evaluator Provider File
  2. Register with the Factory
  3. Add Optional Dependencies
  4. Add Configuration Docs
  5. Add Tests
  6. Add an Example
  7. Add Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit 97fa8d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, toml and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ak Dev New Evaluator Provider loads about 3.4k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,054 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yaalalabs/agent-kernel at commit 97fa8d9, republished under its Apache-2.0 licence (© yaalalabs). 1,054 words, ~3,431 tokens.

Download SKILL.mdSave it as .claude/skills/ak-dev-new-evaluator-provider/SKILL.md (or your agent's skills folder).
name
ak-dev-new-evaluator-provider
description
Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Use this skill when you need to give the test framework's pluggable AKEvaluator interface a new first-party scoring/judge backend addressable by a short config name (e.g. "trulens"), not a one-off bring-your-own evaluator. Covers implementing score-based and LLM-as-judge evaluation, factory registration, configuration, optional dependencies, and testing.
license
Apache-2.0
metadata.author
yaalalabs
metadata.category
developer

Adding a New Test Evaluator Provider

This guide walks through adding a new built-in evaluator provider to Agent Kernel's test framework. Use the existing DeepEval implementation (ak-py/src/agentkernel/test/core/evaluator/deepeval.py) as reference.

Before starting, check whether you actually need this skill: if the evaluator only needs to exist for your own project (not addressable by every AK user via a short built-in name), you don't need any of the steps below — just subclass AKEvaluator anywhere importable and point test-config.yaml's evaluator: at its dotted path. That's the "bring your own evaluator" path described in the user-facing ak-test skill and docs/docs/testing/cli-testing.md; examples/cli/custom-evaluator/ is a complete worked example of it. This skill is only for adding a first-party, in-repo provider that ships with AK and gets its own short type name.

Existing Providers

ProviderShort nameScoring modeLLM-judge modeExtra
DeepEvaldeepevalScorer.quasi_exact_match_score (whole-string, normalised)GEval LLM-as-judge metricagentkernel[test]
OpikopikLevenshteinRatio (fuzzy string similarity)GEval LLM-as-judge metricagentkernel[opik]
JEVjev— (AKMetricNotSupported)TypeSafe Noul (yes/no probability)agentkernel[jev]

Architecture Overview

  • AKEvaluator (ak-py/src/agentkernel/test/core/evaluator/base.py) is the abstract base every evaluator — built-in or bring-your-own — implements. It has exactly two abstract methods:
    • evaluate_by_score(case: AKEvaluationCase) -> AKEvaluationResult — deterministic scoring, no LLM call
    • evaluate_by_llm(case: AKEvaluationCase) -> AKEvaluationResult — LLM-as-judge scoring
  • AKEvaluationCase carries the comparison inputs (user_input, actual, expected, threshold, context, criteria); AKEvaluationResult carries the outcome (score, passed, metric, evaluator, reason, cost, attempts, metadata).
  • Error contract (every evaluator must honor this, not just DeepEval):
    • Raise AKMissingInput when a field the requested metric needs (e.g. case.expected) wasn't supplied.
    • Raise AKMetricNotSupported from whichever of the two methods your backend structurally cannot implement (e.g. a pure LLM-judge service has no offline scoring mode).
    • Raise AKEvaluationError when a configured backend fails to produce a score (missing credentials, transport error, unparseable judge output). Never raise AssertionError and never silently return a 0.0 to stand in for a failure — 0.0 must only ever mean "scored zero", not "couldn't be scored". Test.compare is the only place that decides pass/fail fatality; evaluators only ever set result.passed.
  • Test._resolve_evaluator_class in ak-py/src/agentkernel/test/test.py is the factory. It shares the same pluggable-backend shape as guardrails, sandbox providers, and trace backends (core/util/factory.py's resolve_dotted/require_extra/AKConfigError): an if-per-built-in branch with the SDK import wrapped in require_extra (actionable ImportError naming the pip extra if missing), then a dotted-path bring-your-own fallback for anything else.
  • Evaluator instances are cached per-process on Test._evaluator (keyed by the configured value), guarded by Test._evaluator_lock — construction happens once per distinct evaluator: config value, not once per Test.compare call.

Step-by-Step

1. Create the Evaluator Provider File

Create ak-py/src/agentkernel/test/core/evaluator/<provider>.py. Keep the provider's SDK imports inside this file only — test/core/evaluator/__init__.py and base.py stay pure Python with no optional-dependency imports at module level, so importing the AKEvaluator interface never requires your provider's SDK to be installed.

python
# ak-py/src/agentkernel/test/core/evaluator/<provider>.py
from agentkernel.test.config import AKTestConfig

from .base import AKEvaluationCase, AKEvaluationError, AKEvaluationResult, AKEvaluator, AKMissingInput


class <Provider>AKEvaluator(AKEvaluator):
    def __init__(self, config: AKTestConfig) -> None:
        super().__init__(config)
        # Lazy-init any client/model here only if evaluate_by_score never needs it
        # (mirrors DeepevalAKEvaluator's lazy LiteLLMModel, built only on first evaluate_by_llm call).

    def evaluate_by_score(self, case: AKEvaluationCase) -> AKEvaluationResult:
        if not case.expected:
            raise AKMissingInput("evaluate_by_score requires AKEvaluationCase.expected")
        # Deterministic, offline scoring logic here.
        score = ...  # float
        return AKEvaluationResult(
            metric="<metric_name>",
            evaluator="<provider>",
            score=score,
            passed=score >= case.threshold,
        )

    def evaluate_by_llm(self, case: AKEvaluationCase) -> AKEvaluationResult:
        if not case.expected:
            raise AKMissingInput("evaluate_by_llm requires AKEvaluationCase.expected")
        try:
            score = ...  # call the judge
        except Exception as exc:
            raise AKEvaluationError(f"<provider> llm-based evaluation failed: {exc}") from exc
        return AKEvaluationResult(
            metric="<metric_name>",
            evaluator="<provider>",
            score=score,
            reason=...,  # judge's explanation, if the backend provides one
            passed=score is not None and score >= case.threshold,
        )

If a mode genuinely doesn't apply to your backend (e.g. a provider that is LLM-judge-only), raise AKMetricNotSupported from that method instead of faking a result. Test.compare does not catch it — in fallback mode it propagates out of evaluate_by_score before evaluate_by_llm runs — so document that users of your provider must set the matching mode (e.g. JEV requires mode: llm).

2. Register with the Factory

Add the short name to _BUILTIN_EVALUATORS and a branch in Test._resolve_evaluator_class, both in ak-py/src/agentkernel/test/test.py:

python
_BUILTIN_EVALUATORS = ["deepeval", "opik", "jev", "<provider>"]          # ADD THIS

class Test:
    ...
    @classmethod
    def _resolve_evaluator_class(cls, configured: str) -> type[AKEvaluator]:
        if configured == "deepeval":
            with require_extra("test", "evaluator: deepeval"):
                from .core.evaluator.deepeval import DeepevalAKEvaluator
            return DeepevalAKEvaluator
        if configured == "opik":
            with require_extra("opik", "evaluator: opik"):
                from .core.evaluator.opik import OpikAKEvaluator
            return OpikAKEvaluator
        if configured == "jev":
            with require_extra("jev", "evaluator: jev"):
                from .core.evaluator.jev import JevAKEvaluator
            return JevAKEvaluator
        if configured == "<provider>":                                        # ADD THIS
            with require_extra("<provider>", "evaluator: <provider>"):
                from .core.evaluator.<provider> import <Provider>AKEvaluator
            return <Provider>AKEvaluator
        if "." not in configured:
            raise AKConfigError(
                f"unknown evaluator '{configured}'; expected one of {_BUILTIN_EVALUATORS} or a dotted path to an AKEvaluator subclass"
            )
        return resolve_dotted(configured, base=AKEvaluator)

A dotted evaluator: value (e.g. myorg.evaluators.CustomEvaluator) resolves via resolve_dotted without any factory edit at all — only add an if branch here for a first-party, in-repo provider you want addressable by a short name.

3. Add Optional Dependencies

Add a new extras group to ak-py/pyproject.toml for the provider's SDK — don't fold it into the existing test extra (that one stays DeepEval's, since every test user already needs it for the framework itself). Follow the pattern of the opik extra, the first provider added on top of the original DeepEval-only test extra:

toml
[project.optional-dependencies]
<provider> = [
    "provider-sdk>=x.y.z",
]
Show full SKILL.md (446 more words)Show less
4. Add Configuration Docs

evaluator: in test-config.yaml is already a free-form string on AKTestConfig (built-in short name or dotted path) — no config schema change is needed for a new built-in, since it's just a new value the same field accepts:

yaml
mode: fallback
evaluator: <provider>

If your provider needs extra config fields (e.g. an API key env var name, a judge model override), read them from AKTestConfig the same way DeepevalAKEvaluator reads self._config.llm — don't invent a parallel config path.

5. Add Tests

Add ak-py/tests/test_evaluator_<provider>.py, following the shape of ak-py/tests/test_evaluator_deepeval.py: exercise evaluate_by_score for real (offline, no network) where possible, and mock the judge call in evaluate_by_llm so the suite stays network-free. At minimum cover:

  • evaluate_by_score: exact/mismatch cases, threshold boundary, AKMissingInput when expected is absent
  • evaluate_by_llm: success, failure wrapped as AKEvaluationError, AKMissingInput when expected is absent
  • The factory branch: Test._resolve_evaluator_class("<provider>") resolves to your class, and (if the SDK is optional) the require_extra ImportError path when it's missing — see test_resolve_evaluator_class_deepeval_missing_extra_raises_import_error in ak-py/tests/test_cli_tester.py for the pattern (patching builtins.__import__, since a cached submodule import can otherwise mask the missing dependency).
6. Add an Example

Add examples/cli/<provider>-evaluator/, following the shape of examples/cli/opik-evaluator/ (a minimal agent, a demo_test.py exercising the new evaluator, and a test-config.yaml pointing evaluator: at the new short name). Register it in .github/test-config.yaml's e2e matrix so it runs in CI, the way every other examples/cli/* entry does.

7. Add Documentation

Neither doc page carries a literal "evaluator backend table" — both describe the built-ins in prose next to the score/llm/fallback mode explanations. Update every prose mention that enumerates the built-ins by name, not just one page:

  • docs/docs/core-concepts/configuration.md and docs/docs/testing/cli-testing.md — the evaluator: field description and the score/llm mode explanations.
  • docs/docs/testing/automated-testing.md and docs/docs/testing/overview.md — same prose pattern, duplicated across these pages.
  • docs/docs/agent-skills.md — the skill directory rows for this skill and for ak-dev-testing-conventions.
  • .agents/skills/ak-dev-testing-conventions/SKILL.md — the evaluator config/mode section.
  • ak-py/README.md — the Test Configuration reference (evaluator field) and the test-config walkthrough section.
  • The user-facing ak-test skill (ak-py/src/agentkernel/skills/ak-test/SKILL.md) and its evals/evals.json.
  • Landing page inventories (docs/src/components/*/data.tsx): a tile in the Observability, safety & testing row of IntegrationsMarquee/data.tsx (role Evaluator, href to the automated testing page, logo or react-icons/si glyph), and the provider in the Pluggable Evaluators card's tags and description under the Observe tab in FeatureExplorer/data.tsx. Logo sourcing and the build check are in ak-dev-sync-docs-from-branch, Docs-Site Landing and Features Pages.
  • The features page (docs/src/pages/features.tsx): the approaches entry for Pluggable Evaluators under Testing & Evaluation names every built-in.

Checklist

  • ak-py/src/agentkernel/test/core/evaluator/<provider>.py implementing AKEvaluator
  • Factory registration in Test._resolve_evaluator_class (ak-py/src/agentkernel/test/test.py) and _BUILTIN_EVALUATORS
  • Optional dependency extra in ak-py/pyproject.toml
  • Unit tests in ak-py/tests/test_evaluator_<provider>.py
  • Example in examples/cli/<provider>-evaluator/, registered in .github/test-config.yaml
  • Documentation updated: docs/docs/core-concepts/configuration.md, docs/docs/testing/cli-testing.md, docs/docs/testing/automated-testing.md, docs/docs/testing/overview.md, docs/docs/agent-skills.md, .agents/skills/ak-dev-testing-conventions/SKILL.md, ak-py/README.md, the ak-test skill and its evals/evals.json
  • Landing page inventories: marquee tile (IntegrationsMarquee/data.tsx), Pluggable Evaluators card tags (FeatureExplorer/data.tsx); the Pluggable Evaluators approaches entry in features.tsx

© yaalalabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ak-dev-new-evaluator-provider of yaalalabs/agent-kernel.

Open the folder on GitHubat commit 97fa8d9

Compare with similar skills

Ak Dev New Evaluator Provider next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ak Dev New Evaluator Provider compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ak Dev New Evaluator Provider this skillyaalalabs/agent-kernel192—~3.4kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from yaalalabs/agent-kernel

All 23 skills in this repo
  • Ak Dev Code Quality

    yaalalabs/agent-kernel

    Code quality standards, formatting, Python style rules (classes over script-style functions, configuration-field rules), commit conventions, and PR workflow for Agent Kernel development.

    192 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Ak Dev New Guardrail Provider

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new guardrail provider to Agent Kernel.

    192 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Step-by-step guide for adding a new knowledge base backend to Agent Kernel.

    192 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Ak Dev New Messaging Integration

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new messaging platform integration to Agent Kernel.

    192 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Ak Dev New Multimodal Storage

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new multimodal attachment storage backend to Agent Kernel.

    192 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Ak Dev New Queue Transport

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new queue transport to Agent Kernel's execution pipeline.

    192 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed

Questions about Ak Dev New Evaluator Provider

What does Ak Dev New Evaluator Provider do?

Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Ak Dev New Evaluator Provider is an agent skill from yaalalabs/agent-kernel. Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

When should I use Ak Dev New Evaluator Provider?

Ak Dev New Evaluator Provider fits situations like: tasks that involve LLM evaluation.

How do I install Ak Dev New Evaluator Provider in Claude Code?

Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a claude-code`. Or copy the skill folder (.agents/skills/ak-dev-new-evaluator-provider in yaalalabs/agent-kernel) into .claude/skills/ak-dev-new-evaluator-provider in your project. Claude Code loads it when a task matches its description.

How do I install Ak Dev New Evaluator Provider in Codex?

Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a codex`. Or copy the skill folder (.agents/skills/ak-dev-new-evaluator-provider in yaalalabs/agent-kernel) into .agents/skills/ak-dev-new-evaluator-provider in your project. Codex loads it when a task matches its description.

Can I use Ak Dev New Evaluator Provider in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ak-dev-new-evaluator-provider, .gemini/skills/ak-dev-new-evaluator-provider, .github/skills/ak-dev-new-evaluator-provider and .opencode/skills/ak-dev-new-evaluator-provider in your project.

What does Ak Dev New Evaluator Provider need to run?

SKILL.md names no scripts, command-line tools or credentials: Ak Dev New Evaluator Provider is instructions for the agent only. Our summary lists: Python 3.

Does Ak Dev New Evaluator Provider access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ak Dev New Evaluator Provider safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ak Dev New Evaluator Provider use?

Ak Dev New Evaluator Provider is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ak Dev New Evaluator Provider use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ak Dev New Evaluator Provider?

Skills that share tags, products or a category with Ak Dev New Evaluator Provider: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ak Dev New Evaluator Provider?

yaalalabs (a GitHub organization) maintains it in yaalalabs/agent-kernel, which has 192 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 9, 2026.

Source: yaalalabs/agent-kernel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.